2 papers
cs.AI2026
Bad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts
Guojun Zhu, Xunheng Huang, Peng Yin +3
Reliable agent evaluation is complicated by automatic harness optimization, which repeatedly uses a released benchmark to guide a Proposer that edits prompts, me…
cs.RO2026
FluxVLA Engine: A One-Stop VLA Engineering Platform for Embodied Intelligence
Yinhao Li, Weixin Mao, Zihan Lan +21
Vision-language-action (VLA) models, world-action models (WAMs), and offline reinforcement learning methods are rapidly expanding the design space of embodied policies, yet turning…