2 papers
cs.CV2026
PEEK: Picking Essential frames via Efficient Knowledge distillation
Killian Steunou, Anas Filali Razzouki, Khalil Guetari +3
Video-language models can process only a limited number of frames, making frame selection a key bottleneck for efficient video captioning. Most captioning pipelines still rely on u…
cs.CV2026
Frame Sampling Strategies Matter: A Benchmark for small vision language models
Marija Brkic, Anas Filali Razzouki, Yannis Tevissen +2
Comparing vision language models on videos is particularly complex, as the performances is jointly determined by the model's visual representation capacity and the frame-sampling s…