4 papers · 1 filter
GuideMe: Multi-Domain Task Guidance and Intervention in Streaming Video
Fang Liu, Jinpeng Chen, Ke Xu +7
While multimodal Large Language Models (MLLMs) excel at offline video understanding, an interesting question of how far they are from serving as a real-time procedural coach remain…
X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding
Peiwen Sun, Xudong Lu, Huadai Liu +10
While video streaming understanding has made significant strides, real-world applications, such as live sports broadcasting, autonomous driving, and multi-screen collaboration, inh…
AURA: Always-On Understanding and Real-Time Assistance via Video Streams
Xudong Lu, Yang Bo, Jinpeng Chen +9
Video Large Language Models (VideoLLMs) have achieved strong performance on many video understanding tasks, but most existing systems remain offline and are not well-suited for liv…
Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech
Rui Liu, Shuwei He, Yifan Hu +1
Visual Text-to-Speech (VTTS) aims to take the environmental image as the prompt to synthesize the reverberant speech for the spoken content. The challenge of this task lies in unde…