collaborators

6 papers

cs.CV2026

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding

Wang Chen, Yu Chen, Xiang Wang +3

Frame selection is essential for applying Large Multimodal Models (LMMs) to long videos due to severe frame redundancy and limited context windows. Since the appropriate frame budg…

cs.CV2026

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation

Yuhui Zeng, Wang Chen, Jinfa Huang +5

Existing Large Vision-Language Models (LVLMs) struggle with long-form video understanding due to the quadratic computational cost of visual tokens. While recent efficient methods a…

cs.AI2026

SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models

Tianyu Xie, Jinfa Huang, Yuexiao Ma +11

Omni-modal large language models (OLMs) redefine human-machine interaction by natively integrating audio, vision, and text. However, existing OLM benchmarks remain anchored to stat…

cs.CV2026

Benchmarking Living-Screen-Native GUI Agents on Short-Video Platforms

Jiashu Yao, Heyan Huang, Daiqing Wu +5

GUI agents today assume a static screen, where the world is frozen between two actions. However, real interfaces such as short-video applications violate this assumption, as their…

cs.CV2026

Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding

Wang Chen, Yuhui Zeng, Yongdong Luo +5

Frame selection is crucial due to high frame redundancy and limited context windows when applying Large Vision-Language Models (LVLMs) to long videos. Current methods typically sel…

cs.CV2026

Event-Anchored Frame Selection for Effective Long-Video Understanding

Wang Chen, Yongdong Luo, Yuhui Zeng +5

Massive frame redundancy and limited context window make efficient frame selection crucial for long-video understanding with large vision-language models (LVLMs). Prevailing approa…