3 papers
cs.CV2026
ViterbiPlanNet: Injecting Procedural Knowledge via Differentiable Viterbi for Planning in Instructional Videos
Luigi Seminara, Davide Moltisanti, Antonino Furnari
Procedural planning aims to predict a sequence of actions that transforms an initial visual state into a desired goal, a fundamental ability for intelligent agents operating in com…
cs.CV2024
Continual Learning Improves Zero-Shot Action Recognition
Shreyank N Gowda, Davide Moltisanti, Laura Sevilla-Lara
Zero-shot action recognition requires a strong ability to generalize from pre-training and seen classes to novel unseen classes. Similarly, continual learning aims to develop model…
cs.CV2024
Coarse or Fine? Recognising Action End States without Labels
Davide Moltisanti, Hakan Bilen, Laura Sevilla-Lara +1
We focus on the problem of recognising the end state of an action in an image, which is critical for understanding what action is performed and in which manner. We study this focus…