collaborators

13 papers

cs.CV2026

GuideMe: Multi-Domain Task Guidance and Intervention in Streaming Video

Fang Liu, Jinpeng Chen, Ke Xu +7

While multimodal Large Language Models (MLLMs) excel at offline video understanding, an interesting question of how far they are from serving as a real-time procedural coach remain…

eess.IV2026

See Silhouettes in Motion with Neuromorphic Vision

Pei Zhang, Shijie Lin, Zhou Ge +2

Quasi-bimodal objects, such as text, road signs, and barcodes, play a basic yet vital role in daily visual communication. By boiling these down to clear silhouettes, binarization u…

cs.CV2026

X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding

Peiwen Sun, Xudong Lu, Huadai Liu +10

While video streaming understanding has made significant strides, real-world applications, such as live sports broadcasting, autonomous driving, and multi-screen collaboration, inh…

cs.LG2026

HARD-KV: Head-Adaptive Regularization for Decoding-time KV Compression

Yuxuan Yang, Feiyang Ren, Bowen Zeng +4

Long-context LLM inference faces a fundamental conflict: head-adaptive compression algorithms (e.g., Top- nucleus sampling) offer superior accuracy by dynamically fluctuating me…

cs.AI2026

UI-KOBE: Knowledge-Oriented Behavior Exploration for Lightweight Graph-Guided GUI Agents

Yuxiang Chai, Han Xiao, Xinyu Fu +3

Recent advances in mobile GUI agents have shown strong potential for automating mobile tasks, but most effective systems still depend on large vision-language models for screenshot…

cs.CV2026

OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants

Xudong Lu, Xueying Li, Annan Wang +8

We introduce OmniInteract, a streaming benchmark for real-time omnimodal large language models evaluated through native online inference over audio-visual streams. Unlike offline v…