activity
20242026
collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV2026

SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding

Sara Ghazanfari, Siddharth Garg, Prashanth Krishnamurthy +1

Multimodal Large Language Models (MLLMs) have made rapid progress in single-video understanding, yet their ability to reason across multiple independent video streams remains poorl…

cs.CV2026

BOP-ASK: Object-Interaction Reasoning for Vision-Language Models

Vineet Bhat, Sungsu Kim, Valts Blukis +6

Vision Language Models (VLMs) have achieved impressive performance on spatial reasoning benchmarks, yet these evaluations mask critical weaknesses in understanding object interacti…

cs.CV2026

Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning

Sara Ghazanfari, Francesco Croce, Nicolas Flammarion +3

Recent work has shown that eliciting Large Language Models (LLMs) to generate reasoning traces in natural language before answering the user's request can significantly improve the…

cs.CV2025

EMMA: Efficient Visual Alignment in Multi-Modal LLMs

Sara Ghazanfari, Alexandre Araujo, Prashanth Krishnamurthy +2

Multi-modal Large Language Models (MLLMs) have recently exhibited impressive general-purpose capabilities by leveraging vision foundation models to encode the core concepts of imag…

cs.CV2025

RAZER: Robust Accelerated Zero-Shot 3D Open-Vocabulary Panoptic Reconstruction with Spatio-Temporal Aggregation

Naman Patel, Prashanth Krishnamurthy, Farshad Khorrami

Mapping and understanding complex 3D environments is fundamental to how autonomous systems perceive and interact with the physical world, requiring both precise geometric reconstru…

cs.CV2024

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics

Sara Ghazanfari, Siddharth Garg, Nicolas Flammarion +3

Human perception of similarity across uni- and multimodal inputs is highly complex, making it challenging to develop automated metrics that accurately mimic it. General purpose vis…