13 papers
GuideMe: Multi-Domain Task Guidance and Intervention in Streaming Video
Fang Liu, Jinpeng Chen, Ke Xu +7
While multimodal Large Language Models (MLLMs) excel at offline video understanding, an interesting question of how far they are from serving as a real-time procedural coach remain…
See Silhouettes in Motion with Neuromorphic Vision
Pei Zhang, Shijie Lin, Zhou Ge +2
Quasi-bimodal objects, such as text, road signs, and barcodes, play a basic yet vital role in daily visual communication. By boiling these down to clear silhouettes, binarization u…
X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding
Peiwen Sun, Xudong Lu, Huadai Liu +10
While video streaming understanding has made significant strides, real-world applications, such as live sports broadcasting, autonomous driving, and multi-screen collaboration, inh…
HARD-KV: Head-Adaptive Regularization for Decoding-time KV Compression
Yuxuan Yang, Feiyang Ren, Bowen Zeng +4
Long-context LLM inference faces a fundamental conflict: head-adaptive compression algorithms (e.g., Top- nucleus sampling) offer superior accuracy by dynamically fluctuating me…
UI-KOBE: Knowledge-Oriented Behavior Exploration for Lightweight Graph-Guided GUI Agents
Yuxiang Chai, Han Xiao, Xinyu Fu +3
Recent advances in mobile GUI agents have shown strong potential for automating mobile tasks, but most effective systems still depend on large vision-language models for screenshot…
OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants
Xudong Lu, Xueying Li, Annan Wang +8
We introduce OmniInteract, a streaming benchmark for real-time omnimodal large language models evaluated through native online inference over audio-visual streams. Unlike offline v…