2 papers
cs.CV2025
Chrono: A Simple Blueprint for Representing Time in MLLMs
Hector Rodriguez, Boris Meinardus, Anil Batra +2
The recent success of Large Language Models (LLMs) has prompted the extension to the multimodal domain, developing image-text Multimodal LLMs (MLLMs) and then video-text models. In…
cs.CL2025
Predicting Implicit Arguments in Procedural Video Instructions
Anil Batra, Laura Sevilla-Lara, Marcus Rohrbach +1
Procedural texts help AI enhance reasoning about context and action sequences. Transforming these into Semantic Role Labeling (SRL) improves understanding of individual steps by id…