collaborators

13 papers

cs.CV2026

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding

Wang Chen, Yu Chen, Xiang Wang +3

Frame selection is essential for applying Large Multimodal Models (LMMs) to long videos due to severe frame redundancy and limited context windows. Since the appropriate frame budg…

cs.CV2026

TimePLE: Rethinking Temporal Representation for Video Temporal Grounding

Yuhui Zeng, Xinyu Mao, Xiaokun Liu +4

Video temporal grounding (VTG) aims to localize the continuous video interval described by a natural-language query. However, current VLM-based methods typically produce this inter…

cs.CV2026

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation

Yuhui Zeng, Wang Chen, Jinfa Huang +5

Existing Large Vision-Language Models (LVLMs) struggle with long-form video understanding due to the quadratic computational cost of visual tokens. While recent efficient methods a…

cs.CV2026

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models

Jinsen Su, Yongdong Luo, Yuexiao Ma +4

Existing token compression methods for omnimodal large language models typically rely on one modality to determine what to retain in the other. We show that this assumption often b…

cs.CV2026

SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning

Haoyu Huang, Jinfa Huang, Zhongwei Wan +3

Agentic multimodal large language models (MLLMs) (e.g., OpenAI o3 and Gemini Agentic Vision) achieve remarkable reasoning capabilities through iterative visual tool invocation. How…

cs.CV2026

OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping

Xudong Li, Mengdan Zhang, Peixian Chen +7

Spatial intelligence remains a persistent challenge for Multimodal Large Language Models (MLLMs), as it requires coherent spatial scene representations beyond basic object recognit…