From the 2 of 121 linked papers with an AI index.
121 papers
Masked Visual Actions for Unified World Modeling
Hadi Alzayer, Wenlong Huang, Haonan Chen +8
Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challe…
LOTAPO: Leave-One-Turn Attribution for Self-Generated Process Rewards in Multi-Turn Search Reasoning
Qiang Zhu, Jiajun Wu, Longyi Wang
Reinforcement learning for multi-turn search reasoning typically relies on terminal outcome rewards, which cannot distinguish useful, redundant, and harmful intermediate interactio…
A vision foundation model for single-cell biology via spatial gene cartography
Ridvan Yesiloglu, Sakib Mostafa, James Zou +5
The paper introduces scVision, a vision foundation model that converts single-cell transcriptomic data into images by mapping genes onto a spatial layout, and uses a pretrained vis…
APIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts
Emily Jin, Joy Hsu, Yiqing Xu +3
Long-horizon robot planning requires jointly reasoning over semantic task structure and geometric feasibility. To successfully execute a task, a robot must decompose goals, select…
A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding
Christina Liu, Alan Q. Wang, Joy Hsu +2
Recent tool-use frameworks powered by vision-language models (VLMs) improve image understanding by grounding model predictions with specialized tools. Broadly, these frameworks lev…
Explain Before You Answer: A Survey on Compositional Visual Reasoning
Fucai Ke, Joy Hsu, Zhixi Cai +10
Compositional visual reasoning has emerged as a key research frontier in multimodal AI, aiming to endow machines with the human-like ability to decompose visual scenes, ground inte…