Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Allocate Before You Embed: Adaptive Visual Input Allocation for Video Embeddings
Song Jin, Zhongtao Jiang, Chenglei Shen +5
Large-scale video retrieval requires embedding models to encode long and diverse videos under tight visual-input and inference budgets. Existing methods typically sample a small, f…
cs.CV2026
ResAdapt: Adaptive Resolution for Efficient Multimodal Reasoning
Huanxuan Liao, Zhongtao Jiang, Yupu Hao +6
Multimodal Large Language Models (MLLMs) achieve stronger visual understanding by scaling input fidelity, yet the resulting visual token growth makes jointly sustaining high spatia…