5 papers
Predicting Performance of Symbolic and Prompt Programs with Examples
Chengqi Zheng, Keya Hu, Shuzhi Liu +3
LLM prompting is widely used for naturally stated tasks, yet it is unreliable it may succeed on a few test cases but fail at deployment time. We study performance prediction: given…
System Design for Maintaining Internal State Consistency in Long-Horizon Robotic Tabletop Games
Guangyu Zhao, Ceyao Zhang, Chengdong Ma +16
Long-horizon tabletop games pose a distinct systems challenge for robotics: small perceptual or execution errors can invalidate accumulated task state, propagate across decision-ma…
OmniScientist: Toward a Co-evolving Ecosystem of Human and AI Scientists
Chenyang Shao, Dehao Huang, Yu Li +18
With the rapid development of Large Language Models (LLMs), AI agents have demonstrated increasing proficiency in scientific tasks, ranging from hypothesis generation and experimen…
When Robots Should Say "I Don't Know": Benchmarking Abstention in Embodied Question Answering
Tao Wu, Chuhao Zhou, Guangyu Zhao +3
Embodied Question Answering (EQA) requires an agent to interpret language, perceive its environment, and navigate within 3D scenes to produce responses. Existing EQA benchmarks ass…
NoisyEQA: Benchmarking Embodied Question Answering Against Noisy Queries
Tao Wu, Chuhao Zhou, Yen Heng Wong +2
The rapid advancement of Vision-Language Models (VLMs) has significantly advanced the development of Embodied Question Answering (EQA), enhancing agents' abilities in language unde…