Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data
Valentin Gabeff, Baptiste Maquignaz, Jennifer Shan +5
Automatically retrieving videos from large camera-trap datasets remains challenging. Text-to-Video retrieval (TVR) methods based on large video-language models (VLMs) have potentia…
cs.CV2025
VinaBench: Benchmark for Faithful and Consistent Visual Narratives
Silin Gao, Sheryl Mathew, Li Mi +6
Visual narrative generation transforms textual narratives into sequences of images illustrating the content of the text. However, generating visual narratives that are faithful to…