collaborators

5 papers

cs.LG2026

Calibrated Sampling-Free Uncertainty Estimation in Bayesian Deep Learning

Tobias Jan Wieczorek, Leon de Andrade, Thomas Möllenhoff +1

Modern deep learning models remain notoriously prone to overconfidence, limiting their reliability in high-stakes applications. Bayesian methods aim to counter this by learning a d…

cs.CV2026

ReCap: Lightweight Referential Grounding for Coherent Story Visualization

Aditya Arora, Akshita Gupta, Pau Rodriguez +1

Story Visualization aims to generate a sequence of images that faithfully depicts a textual narrative that preserve character identity, spatial configuration, and stylistic coheren…

cs.CV2026

Evaluating the Impact of Post-Training Quantization on Reliable VQA with Multimodal LLMs

Paul Jonas Kurz, Tobias Jan Wieczorek, Mohamed A. Abdelsalam +2

Multimodal Large Language Models (MLLM) are increasingly deployed in domains where both reliability and efficiency are critical. However, current models remain overconfident, produ…

cs.CV2025

Chrono: A Simple Blueprint for Representing Time in MLLMs

Hector Rodriguez, Boris Meinardus, Anil Batra +2

The recent success of Large Language Models (LLMs) has prompted the extension to the multimodal domain, developing image-text Multimodal LLMs (MLLMs) and then video-text models. In…

cs.CL2025

Predicting Implicit Arguments in Procedural Video Instructions

Anil Batra, Laura Sevilla-Lara, Marcus Rohrbach +1

Procedural texts help AI enhance reasoning about context and action sequences. Transforming these into Semantic Role Labeling (SRL) improves understanding of individual steps by id…