4 papers · 1 filter
ReCap: Lightweight Referential Grounding for Coherent Story Visualization
Aditya Arora, Akshita Gupta, Pau Rodriguez +1
Story Visualization aims to generate a sequence of images that faithfully depicts a textual narrative that preserve character identity, spatial configuration, and stylistic coheren…
Evaluating the Impact of Post-Training Quantization on Reliable VQA with Multimodal LLMs
Paul Jonas Kurz, Tobias Jan Wieczorek, Mohamed A. Abdelsalam +2
Multimodal Large Language Models (MLLM) are increasingly deployed in domains where both reliability and efficiency are critical. However, current models remain overconfident, produ…
Chrono: A Simple Blueprint for Representing Time in MLLMs
Hector Rodriguez, Boris Meinardus, Anil Batra +2
The recent success of Large Language Models (LLMs) has prompted the extension to the multimodal domain, developing image-text Multimodal LLMs (MLLMs) and then video-text models. In…
Efficient Pre-training for Localized Instruction Generation of Videos
Anil Batra, Davide Moltisanti, Laura Sevilla-Lara +2
Procedural videos, exemplified by recipe demonstrations, are instrumental in conveying step-by-step instructions. However, understanding such videos is challenging as it involves t…