collaborators

5 papers

cs.AI2026

Allocation Before Ranking: Decoupled Token Compression for OmniLLMs

Zhenghui Guo, Yilin Yang, Yuanbin Man +5

Token compression in OmniLLMs is typically posed as a single saliency-ranking problem: score each multimodal token, keep the top-K. We argue this abstraction is mis-specified. The…

cs.GR2026

Turbo4DGen: Ultra-Fast Acceleration for 4D Generation

Yuanbin Man, Ying Huang, Zhile Ren +1

4D generation, or dynamic 3D content generation, integrates spatial, temporal, and view dimensions to model realistic dynamic scenes, playing a foundational role in advancing world…

cs.CV2026

Event-VStream: Event-Driven Real-Time Understanding for Long Video Streams

Zhenghui Guo, Yuanbin Man, Junyuan Sheng +8

Real-time understanding of long video streams remains challenging for multimodal large language models (VLMs) due to redundant frame processing and rapid forgetting of past context…

cs.CV2025

AdaRing: Towards Ultra-Light Vision-Language Adaptation via Cross-Layer Tensor Ring Decomposition

Ying Huang, Yuanbin Man, Wenqi Jia +3

Adapter-based fine-tuning has gained remarkable attention in adapting large pre-trained vision language models (VLMs) for a wide range of downstream tasks efficiently. In this para…

cs.CV2025

AdaCM: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction

Yuanbin Man, Ying Huang, Chengming Zhang +3

The advancements in large language models (LLMs) have propelled the improvement of video understanding tasks by incorporating LLMs with visual models. However, most existing LLM-ba…