6 papers
GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation
Kaichen Zhou, Yuzhen Chen, Fangneng Zhan +8
Video world models can generate realistic futures from a single instruction, but they often fail to track the same physical points consistently across time. As a result, the genera…
SVR-MAD: A Bayesian-Inspired Framework for Posterior-Guided Multi-Agent Debate
Weifan Jiang, Rana Shahout, Minghao Li +4
Multi-Agent Debate (MAD) improves LLM-agent accuracy but suffers from rapid context growth, limiting scalability in larger multi-agent settings. Existing methods prune low-utility…
Plan First, Diffuse Later: Extrinsic Graph Guidance for Long-Horizon Diffusion Planning
Yaniv Hassidof, Adir Morgan, Yilun Du +1
Compositional diffusion models offer a promising route to long-horizon planning by denoising multiple overlapping sub-trajectories while ensuring that together they constitute a gl…
D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting
Tianyu Wu, Yu Yao, Zhenting Qi +7
Speculative decoding accelerates LLM inference by having a small drafter propose tokens that a larger target model verifies in parallel. Recent diffusion-based parallel drafters su…
Anomalies by Synthesis: Anomaly Detection using Generative Diffusion Models for Off-Road Navigation
Siddharth Ancha, Sunshine Jiang, Travis Manderson +4
In order to navigate safely and reliably in off-road and unstructured environments, robots must detect anomalies that are out-of-distribution (OOD) with respect to the training dat…
Inference-Time Policy Steering through Human Interactions
Yanwei Wang, Lirui Wang, Yilun Du +6
Generative policies trained with human demonstrations can autonomously accomplish multimodal, long-horizon tasks. However, during inference, humans are often removed from the polic…