29 papers
Trust Region Inverse Reinforcement Learning: Explicit Dual Ascent using Local Policy Updates
Anish Diwan, Davide Tateo, Christopher E. Mower +3
Inverse reinforcement learning (IRL) is typically formulated as maximizing entropy subject to matching the distribution of expert trajectories. Classical (dual-ascent) IRL guarante…
HapTile: A Haptic-Informed Vision-Tactile-Language-Action Dataset for Contact-Rich Imitation Learning
Amirhosein Alian, Yongqiang Zhao, Shiyi Gu +5
Despite the importance of tactile sensing for reliable manipulation, most existing Vision-Language-Action (VLA) datasets remain vision-only, and those that do incorporate tactile i…
The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling
Tu Nguyen, Matthieu Zimmer, Rasul Tutunov +2
A recurring pattern in "reasoning without training" is that base LLMs already assign non-trivial probability mass to correct multi-step solutions; the bottleneck is locating these…
EmbodimentSemantic: A Spatial Scene-Graph Dataset and Benchmark for Vision-Language Models on Embodied Manipulation Trajectories
Hassan Jaber, Refinath S N, Luca Cagliero +2
Spatial grounding remains a key limitation of vision-language-action (VLA) systems for robotic manipulation. While current models can recognize objects and follow language instruct…
Robots Need More than VLA and World Models
Elis Karcini, Faisal Mehrban, Quang Nguyen +6
Generalist robot intelligence is often framed as a policy-scaling problem: collect more robot demonstrations, train larger Vision-Language-Action (VLA) models, and expect broader g…
Risk-Controlled Lean-as-Judge for Natural-Language Mathematical Reasoning
Pauline Bourigault, Xiaotong Ji, Matthieu Zimmer +2
Lean is increasingly used to judge natural-language mathematical answers, but its signal is partial: many answers never formalize, and a failed proof may reflect an ill-typed state…