15 papers
Align-RAG: Alignment Is All You Need for TSFM In-Context Learning
Mohammad Asadi, Soheil Hor, Bardiya Akhbari +6
Retrieval-augmented forecasting promises to adapt frozen Time Series Foundation Models (TSFMs) to new domains without fine-tuning, but recent methods typically rely on learned fusi…
Ordered Action Tokens for Visuomotor Policy Learning
Chaoqi Liu, Yue Zhao, Haonan Chen +4
Action tokenization maps continuous robot action chunks to discrete tokens and has become an important interface for modern visuomotor policies. Existing approaches either rely on…
A vision foundation model for single-cell biology via spatial gene cartography
Ridvan Yesiloglu, Sakib Mostafa, James Zou +5
The paper introduces scVision, a vision foundation model that converts single-cell transcriptomic data into images by mapping genes onto a spatial layout, and uses a pretrained vis…
A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding
Christina Liu, Alan Q. Wang, Joy Hsu +2
Recent tool-use frameworks powered by vision-language models (VLMs) improve image understanding by grounding model predictions with specialized tools. Broadly, these frameworks lev…
HumanScore: Benchmarking Human Motions in Generated Videos
Yusu Fang, Tiange Xiang, Tian Tan +4
Recent advances in model architectures, compute, and data scale have driven rapid progress in video generation, producing increasingly realistic content. Yet, no prior method syste…
ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body
Juze Zhang, Changan Chen, Xin Chen +5
Human communication is inherently multimodal and social: words, prosody, and body language jointly carry intent. Yet most prior systems model human behavior as a translation task c…