7 papers
FORGE: Towards Functional Tool-Use Generalization via Keypoint Trajectory Reasoning
Chuhao Zhou, Liquan Wang, Shuxin Cao +5
While humans readily repurpose a book, a stone, or a shoe to drive a nail, robots trained on specific tools fail to transfer the same function to novel ones -- a gap we formalize a…
Hierarchical Policy Learning via Spectral Decomposition
Shuxin Cao, Liquan Wang, Walker Byrnes +3
In this paper, we identify a semantic decomposition in robot action sequences, separating task-level motion intent from execution-level refinements. By analyzing actions in the spe…
MosaicMem: Hybrid Spatial Memory for Controllable Video World Models
Wei Yu, Runjia Qian, Yumeng Li +8
Video diffusion models are moving beyond short, plausible clips toward world simulators that must remain consistent under camera motion, revisits, and intervention. Yet spatial mem…
OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model
Ishika Singh, Ankit Goyal, Stan Birchfield +3
We introduce OG-VLA, a novel architecture and learning framework that combines the generalization strengths of Vision Language Action models (VLAs) with the robustness of 3D-aware…
TopoCut: Learning Multi-Step Cutting with Spectral Rewards and Discrete Diffusion Policies
Liquan Wang, Jiangjie Bian, Eric Heiden +1
Robotic manipulation tasks involving cutting deformable objects remain challenging due to complex topological behaviors, difficulties in perceiving dense object states, and the lac…
RoCoDA: Counterfactual Data Augmentation for Data-Efficient Robot Learning from Demonstrations
Ezra Ameperosa, Jeremy A. Collins, Mrinal Jain +1
Imitation learning in robotics faces significant challenges in generalization due to the complexity of robotic environments and the high cost of data collection. We introduce RoCoD…