2 papers
cs.IR2026
Skim and Skip: Hierarchical Adaptive Inference for Efficient Multimodal Retrieval
Meng Gao, Yizhen Zhang, Yang Ding +10
Universal multimodal retrieval (UMR) increasingly adopts multimodal large language models (MLLMs) as unified embedding backbones, but their strong retrieval performance comes at su…
cs.CV2026
GeoTrace: Geometry-Aware Trajectory Token Compression for Video Large Language Models
Guohuan Xie, Mengqi Lei, Chuan Shi +3
Although Video Large Language Models (Video LLMs) have shown strong performance in video understanding, their efficiency is still limited by the large number of visual tokens. Exis…