works on

From the 1 of 45 linked papers with an AI index.

activity
20242026
most citedOn the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective

1 citations · 1 across the 8 of their papers we have counts for

collaborators
Showing cs.CVShow all

11 papers · 1 filter

cs.CV2026

Leveraging Latent Visual Reasoning in Silence

Dongyao Zhu, Zhen Wang, Xi Xiao +7

Latent visual reasoning involves visual evidence more directly in multimodal reasoning by inserting continuous latent tokens before textual generation. However, the necessity of th…

cs.CV2026

RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics

Chan Hee Song, Valts Blukis, Jonathan Tremblay +3

Spatial understanding is a crucial capability that enables robots to perceive their surroundings, reason about their environment, and interact with it meaningfully. In modern robot…

cs.CV2025

Interpretable and Testable Vision Features via Sparse Autoencoders

Samuel Stevens, Wei-Lun Chao, Tanya Berger-Wolf +1

To truly understand vision models, we must not only interpret their learned features but also validate these interpretations through controlled experiments. While earlier work offe…

cs.CV2025

BioCLIP 2: Emergent Properties from Scaling Hierarchical Contrastive Learning

Jianyang Gu, Samuel Stevens, Elizabeth G Campolongo +13

Foundation models trained at scale exhibit remarkable emergent behaviors, learning new capabilities beyond their initial training objectives. We find such emergent behaviors in bio…

cs.CV2025

Prompt-CAM: Making Vision Transformers Interpretable for Fine-Grained Analysis

Arpita Chowdhury, Dipanjyoti Paul, Zheda Mai +10

We present a simple approach to make pre-trained Vision Transformers (ViTs) interpretable for fine-grained analysis, aiming to identify and localize the traits that distinguish vis…

cs.CV2025

Finer-CAM: Spotting the Difference Reveals Finer Details for Visual Explanation

Ziheng Zhang, Jianyang Gu, Arpita Chowdhury +5

Class activation map (CAM) has been widely used to highlight image regions that contribute to class predictions. Despite its simplicity and computational efficiency, CAM often stru…