Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs
Killian Steunou, Yannis Tevissen, Mounîm A. El Yacoubi
Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video representations with pretrained large language models and conditio…
cs.CV2026
PEEK: Picking Essential frames via Efficient Knowledge distillation
Killian Steunou, Anas Filali Razzouki, Khalil Guetari +2
Video-language models can process only a limited number of frames, making frame selection a key bottleneck for efficient video captioning. Most captioning pipelines still rely on u…