5 papers
SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding
Sara Ghazanfari, Siddharth Garg, Prashanth Krishnamurthy +1
Multimodal Large Language Models (MLLMs) have made rapid progress in single-video understanding, yet their ability to reason across multiple independent video streams remains poorl…
Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning
Sara Ghazanfari, Francesco Croce, Nicolas Flammarion +3
Recent work has shown that eliciting Large Language Models (LLMs) to generate reasoning traces in natural language before answering the user's request can significantly improve the…
EMMA: Efficient Visual Alignment in Multi-Modal LLMs
Sara Ghazanfari, Alexandre Araujo, Prashanth Krishnamurthy +2
Multi-modal Large Language Models (MLLMs) have recently exhibited impressive general-purpose capabilities by leveraging vision foundation models to encode the core concepts of imag…
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics
Sara Ghazanfari, Siddharth Garg, Nicolas Flammarion +3
Human perception of similarity across uni- and multimodal inputs is highly complex, making it challenging to develop automated metrics that accurately mimic it. General purpose vis…
Out-of-Distribution Detection with Overlap Index
Hao Fu, Prashanth Krishnamurthy, Siddharth Garg +1
Out-of-distribution (OOD) detection is crucial for the deployment of machine learning models in the open world. While existing OOD detectors are effective in identifying OOD sample…