13 papers
One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding
Wang Chen, Yu Chen, Xiang Wang +3
Frame selection is essential for applying Large Multimodal Models (LMMs) to long videos due to severe frame redundancy and limited context windows. Since the appropriate frame budg…
TimePLE: Rethinking Temporal Representation for Video Temporal Grounding
Yuhui Zeng, Xinyu Mao, Xiaokun Liu +4
Video temporal grounding (VTG) aims to localize the continuous video interval described by a natural-language query. However, current VLM-based methods typically produce this inter…
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation
Yuhui Zeng, Wang Chen, Jinfa Huang +5
Existing Large Vision-Language Models (LVLMs) struggle with long-form video understanding due to the quadratic computational cost of visual tokens. While recent efficient methods a…
OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models
Jinsen Su, Yongdong Luo, Yuexiao Ma +4
Existing token compression methods for omnimodal large language models typically rely on one modality to determine what to retain in the other. We show that this assumption often b…
SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning
Haoyu Huang, Jinfa Huang, Zhongwei Wan +3
Agentic multimodal large language models (MLLMs) (e.g., OpenAI o3 and Gemini Agentic Vision) achieve remarkable reasoning capabilities through iterative visual tool invocation. How…
OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping
Xudong Li, Mengdan Zhang, Peixian Chen +7
Spatial intelligence remains a persistent challenge for Multimodal Large Language Models (MLLMs), as it requires coherent spatial scene representations beyond basic object recognit…