collaborators

5 papers

cs.CV2026

BioCAP: Exploiting Synthetic Captions Beyond Labels in Biological Foundation Models

Ziheng Zhang, Xinyue Ma, Arpita Chowdhury +9

This work investigates descriptive captions as an additional source of supervision for biological multimodal foundation models. Images and captions can be viewed as complementary s…

cs.CV2026

Towards Artwork Explanation in Large-scale Vision Language Models

Kazuki Hayashi, Yusuke Sakai, Hidetaka Kamigaito +2

Large-scale Vision-Language Models (LVLMs) output text from images and instructions, demonstrating capabilities in text generation and comprehension. However, it has not been clari…

cs.CV2025

Towards Open-Ended Visual Scientific Discovery with Sparse Autoencoders

Samuel Stevens, Jacob Beattie, Tanya Berger-Wolf +1

Scientific archives now contain hundreds of petabytes of data across genomics, ecology, climate, and molecular biology that could reveal undiscovered patterns if systematically ana…

cs.CV2025

Interpretable and Testable Vision Features via Sparse Autoencoders

Samuel Stevens, Wei-Lun Chao, Tanya Berger-Wolf +1

To truly understand vision models, we must not only interpret their learned features but also validate these interpretations through controlled experiments. While earlier work offe…

cs.CV2025

BioCLIP 2: Emergent Properties from Scaling Hierarchical Contrastive Learning

Jianyang Gu, Samuel Stevens, Elizabeth G Campolongo +13

Foundation models trained at scale exhibit remarkable emergent behaviors, learning new capabilities beyond their initial training objectives. We find such emergent behaviors in bio…