4 papers
Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence
Ying Chen, Weizhen Li, Zhe Hu +7
Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each action reshapes the scene and physical state…
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack
Siyi Li, Chunyu Sun, Jiahao Zhang +6
Evaluating a Physical AI stack spans operators that differ by more than three orders of magnitude -- from a single foundation-model decoding step to thousands of physics ticks of w…
Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents
Ying Chen, Lihuang Fang, Rui Jiang +4
Standard embodied evaluations do not independently score whether an agent correctly commits to task completion at episode closure, a capacity we call terminal commitment. Behaviora…
BLADE: Benchmarking Language Model Agents for Data-Driven Science
Ken Gu, Ruoxi Shang, Ruien Jiang +13
Data-driven scientific discovery requires the iterative integration of scientific domain knowledge, statistical expertise, and an understanding of data semantics to make nuanced an…