collaborators

6 papers

cs.CV2026

Same or Not? Enhancing Visual Perception in Vision-Language Models

Damiano Marsili, Aditya Mehta, Ryan Y. Lin +1

Vision-language models (VLMs) excel at broad visual understanding but remain coarse-grained, exhibit visual biases, and miss subtle visual details. Existing training corpora reinfo…

cs.CV2025

No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers

Damiano Marsili, Georgia Gkioxari

Visual reasoning is challenging, requiring both precise object grounding and understanding complex spatial relationships. Existing methods fall into two camps: language-only chain-…

cs.AI2025

NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models

Mouadh Yagoubi, Yasser Dahou, Billel Mokeddem +12

Existing benchmarks have proven effective for assessing the performance of fully trained large language models. However, we find striking differences in the early training stages o…

cs.CV2025

Visual Agentic AI for Spatial Reasoning with a Dynamic API

Damiano Marsili, Rohun Agrawal, Yisong Yue +1

Visual reasoning -- the ability to interpret the visual world -- is crucial for embodied agents that operate within three-dimensional scenes. Progress in AI has led to vision and l…

cs.CV2025

Find Any Part in 3D

Ziqi Ma, Yisong Yue, Georgia Gkioxari

Why don't we have foundation models in 3D yet? A key limitation is data scarcity. For 3D object part segmentation, existing datasets are small in size and lack diversity. We show t…

cs.LG2025

TOTEM: TOkenized Time Series EMbeddings for General Time Series Analysis

Sabera Talukder, Yisong Yue, Georgia Gkioxari

This work studies the problem of time series analysis with generalist (or foundation) models, which are models trained across many data domains. Drawing inspiration from the widesp…