11 papers
PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment
Yuyang Liu, Yanqing Shen, Ruike Chen +22
Fine-grained robotic evaluation matters for understanding embodied models, going beyond binary success rates and rule-based process scores. We present PRM-as-a-Judge 1.5, a toolkit…
SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation
Jingkai Wang, Zihan Tang, Gu Zhang +7
Vision-language-action policies rely on large multimodal backbones to jointly perform perception, language conditioning, and action generation at every control step. Much of this c…
SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs
Yadi Cao, Sicheng Lai, Jiahe Huang +12
Evaluating LLM agents for scientific tasks has focused on token costs while ignoring tool-use costs like simulation time and experimental resources. As a result, metrics like pass@…
General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling
Huaihai Lyu, Chaofan Chen, Mingyu Cao +2
Achieving robust generalization from limited data is a central challenge in embodied intelligence. Prevailing methods fail by regressing absolute coordinates, which violates the pr…
RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation
Shihan Wu, Xuecheng Liu, Shaoxuan Xie +81
Despite the critical role of bimanual manipulation in endowing robots with human-like dexterity, large-scale and diverse datasets remain scarce due to the significant hardware hete…
PRM-as-a-Judge: A Dense Evaluation Paradigm for Fine-Grained Robotic Auditing
Yuheng Ji, Yuyang Liu, Huajie Tan +15
Current robotic evaluation is still largely dominated by binary success rates, which collapse rich execution processes into a single outcome and obscure critical qualities such as…