6 papers
GaussianDream++: Efficient 3D Gaussian World Modeling for Robotic Manipulation
Yuqing Jiang, Zijian Zhang, Weitao Zhou +8
Vision-Language-Action (VLA) policies have advanced language-conditioned robotic manipulation, yet action-imitation objectives provide only weak supervision for metric 3D structure…
Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning
Jiawei Wang, Ke Rui, Yushen Zuo +2
JEPA-style latent world models can use Euclidean distance to a goal latent as the cost for model-predictive control (MPC). Strong decoding of task variables, however, does not guar…
SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation
Bowen Jing, Mingxin Wang, Ruiyang Hao +15
Physical interaction quality is central to deformable-object manipulation, yet most benchmarks evaluate task success alone. A policy may complete the task while allowing slip or ca…
SoftVTBench: A Safety-Aware Visuo-Tactile Benchmark for Physically Constrained Robotic Manipulation of Deformable Objects (Early Version)
Bowen Jing, Mingxin Wang, Ruiyang Hao +15
Deformable object manipulation poses challenges beyond task completion: successful execution must also maintain safe physical interaction, holding the object stably without slip or…
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
Simple AI, :, Yuteng Wei +16
Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scalable. Real-robot teleoperation is accurate but costly to scale;…
Diagnosing Semantic Handoff Failures in Agent-Orchestrated Vision-Language-Action Skill Composition
Ke Rui, Yushen Zuo, Jiawei Wang +4
Long-horizon household tasks require robots to compose many language-conditioned skills, yet the boundary between consecutive skills is rarely explicit. A skill may satisfy its own…