8 papers
SABRE: Scalable and Automated Benchmarking of VLMs under Stress
Zixuan Lan, Luzhe Sun, Matthew R. Walter +1
Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making weaknesses hard to identify. Building stress tests is costly: samples must satisf…
Flow-based Policy Adaptation without Policy Updates
Luzhe Sun, Jingtian Ji, Haoran Chen +2
Leveraging prior knowledge from pretrained policies, foundation models, or human operators offers an efficient alternative to learning robot skills from scratch. However, these age…
Robotic Manipulation is Vision-to-Geometry Mapping: Vision-Geometry Backbones over Language and Video Models
Zijian Song, Qichang Li, Jiawei Zhou +4
At its core, robotic manipulation is a problem of vision-to-geometry mapping (). Physical actions are fundamentally defined by geometric properties like 3D posi…
HintMR: Eliciting Stronger Mathematical Reasoning in Small Language Models
Jawad Hossain, Xiangyu Guo, Jiawei Zhou +1
Small language models (SLMs) often struggle with complex mathematical reasoning due to limited capacity to maintain long chains of intermediate steps and to recover from early erro…
Self-Improvement of Large Language Models: A Technical Overview and Future Outlook
Haoyan Yang, Mario Xerri, Solha Park +4
As large language models (LLMs) continue to advance, improving them solely through human supervision is becoming increasingly costly and limited in scalability. As models approach…
HALP: Detecting Hallucinations in Vision-Language Models without Generating a Single Token
Sai Akhil Kogilathota, Sripadha Vallabha E G, Luzhe Sun +1
Hallucinations remain a persistent challenge for vision-language models (VLMs), which often describe nonexistent objects or fabricate facts. Existing detection methods typically op…