1 paper
Yulin Zou, Yan Chen, Wenyan Chen +5
Video streaming analytics is a crucial workload for vision-language model serving, but the high cost of multimodal inference limits scalability. Prior systems reduce inference cost…