2 papers
cs.AI2026
HiL-Bench (Human-in-Loop Benchmark): Do Agents Know When to Ask for Help?
Tu Trinh, Mohamed Elfeki, Guangze Luo +9
Frontier coding agents solve complex tasks when given complete context but collapse when specifications are incomplete or ambiguous. The bottleneck is not raw capability, but judgm…
cs.LG2026
: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities
Physical Intelligence, Bo Ai, Ali Amin +85
We present a new robotic foundation model, called , that can enable strong out-of-the-box performance in a wide range of scenarios. can follow diverse language…