2 papers
cs.CV2025
Training-free Online Video Step Grounding
Luca Zanella, Massimiliano Mancini, Yiming Wang +2
Given a task and a set of steps composing it, Video Step Grounding (VSG) aims to detect which steps are performed in a video. Standard approaches for this task require a labeled tr…
cs.CV2025
Can Text-to-Video Generation help Video-Language Alignment?
Luca Zanella, Massimiliano Mancini, Willi Menapace +3
Recent video-language alignment models are trained on sets of videos, each with an associated positive caption and a negative caption generated by large language models. A problem…