activity
20242026
collaborators

6 papers

cs.LG2026

Abra: Scaling Diffusion Image Training

Kyle Chickering, Wei-An Lin, Swayam Bhanded +5

Compute-optimal scaling laws guide the training of frontier language models yet remain largely unexplored for visual generation. We present a systematic scaling law study for text-…

cs.CV2026

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding

Arsha Nagrani, Jasper Uijilings, Shyamal Buch +6

Video reasoning models are a core component of egocentric and embodied agents. However, standard benchmarks for assessing models provide only evaluation of the output (e.g. the ans…

cs.CV2025

OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models

Monika Wysoczańska, Shyamal Buch, Anurag Arnab +1

Large vision-language models (VLMs) often struggle to generate long and factual captions. However, traditional measures for hallucination and factuality are not well suited for eva…

cs.LG2025

MINERVA: Evaluating Complex Video Reasoning

Arsha Nagrani, Sachit Menon, Ahmet Iscen +9

Multimodal LLMs are turning their focus to video benchmarks, however most video benchmarks only provide outcome supervision, with no intermediate or interpretable reasoning steps.…

cs.CV2025

MoReVQA: Exploring Modular Reasoning Models for Video Question Answering

Juhong Min, Shyamal Buch, Arsha Nagrani +2

This paper addresses the task of video question answering (videoQA) via a decomposed multi-stage, modular reasoning framework. Previous modular methods have shown promise with a si…

cs.CV2024

Streaming Detection of Queried Event Start

Cristobal Eyzaguirre, Eric Tang, Shyamal Buch +3

Robotics, autonomous driving, augmented reality, and many embodied computer vision applications must quickly react to user-defined events unfolding in real time. We address this se…