2 papers
cs.CV2026
EvoGround: Self-Evolving Video Agents for Video Temporal Grounding
Minjoon Jung, Byoung-Tak Zhang, Lorenzo Torresani
Video temporal grounding (VTG) takes an untrimmed video and a natural-language query as input and localizes the temporal moment that best matches the query. Existing methods rely o…
cs.CV2024
Semantic Compositions Enhance Vision-Language Contrastive Learning
Maxwell Aladago, Lorenzo Torresani, Soroush Vosoughi
In the field of vision-language contrastive learning, models such as CLIP capitalize on matched image-caption pairs as positive examples and leverage within-batch non-matching pair…