2 papers
cs.CV2026
A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding
Christina Liu, Alan Q. Wang, Joy Hsu +2
Recent tool-use frameworks powered by vision-language models (VLMs) improve image understanding by grounding model predictions with specialized tools. Broadly, these frameworks lev…
cs.CV2024
Learning Keypoints for Multi-Agent Behavior Analysis using Self-Supervision
Daniel Khalil, Christina Liu, Pietro Perona +2
The study of social interactions and collective behaviors through multi-agent video analysis is crucial in biology. While self-supervised keypoint discovery has emerged as a promis…