10 papers
APIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts
Emily Jin, Joy Hsu, Yiqing Xu +3
Long-horizon robot planning requires jointly reasoning over semantic task structure and geometric feasibility. To successfully execute a task, a robot must decompose goals, select…
A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding
Christina Liu, Alan Q. Wang, Joy Hsu +2
Recent tool-use frameworks powered by vision-language models (VLMs) improve image understanding by grounding model predictions with specialized tools. Broadly, these frameworks lev…
Explain Before You Answer: A Survey on Compositional Visual Reasoning
Fucai Ke, Joy Hsu, Zhixi Cai +10
Compositional visual reasoning has emerged as a key research frontier in multimodal AI, aiming to endow machines with the human-like ability to decompose visual scenes, ground inte…
SAW-Bench: Learning Situated Awareness in the Real World
Chuhan Li, Rilyn Han, Joy Hsu +5
A core aspect of human perception is situated awareness, the ability to relate ourselves to the surrounding physical environment and reason over possible actions in context. Howeve…
A Matter of Interest: Understanding Interestingness of Math Problems in Humans and Language Models
Shubhra Mishra, Yuka Machino, Gabriel Poesia +9
The evolution of mathematics is shaped importantly by interestingness: researchers choose which problems to pursue, and students choose which problems to engage with, based on expe…
Neuro-Symbolic Decoding of Neural Activity
Yanchen Wang, Joy Hsu, Ehsan Adeli +1
We propose NEURONA, a neuro-symbolic framework for fMRI decoding and concept grounding in neural activity. Leveraging image- and video-based fMRI question-answering datasets, NEURO…