11 papers
Foresight Residual RL for Long-Horizon Robot Manipulation with Vision-Language-Action Models
Yuhan Liu, Xinyu Zhang, Litao Liu +1
Vision-Language-Action (VLA) policies offer strong general-purpose manipulation priors, but often fail on tight-tolerance, contact-rich assembly due to long-horizon credit assignme…
Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing Gloves
Xinyu Zhang, Ziyi Kou, Chuan Qin +7
Understanding hand-object interaction (HOI) is fundamental to computer vision, robotics, and AR/VR. However, conventional hand videos often lack essential physical information such…
Learning Visual Feature-Based World Models via Residual Latent Action
Xinyu Zhang, Zhengtong Xu, Yutian Tao +3
World models predict future transitions from observations and actions. Existing works predominantly focus on image generation only. Visual feature-based world models, on the other…
Bellman Diffusion Models
Liam Schramm, Abdeslam Boularias
Diffusion models have seen tremendous success as generative architectures. Recently, they have been shown to be effective at modelling policies for offline reinforcement learning a…
Motion Blender Gaussian Splatting for Dynamic Scene Reconstruction
Xinyu Zhang, Haonan Chang, Yuhan Liu +1
Gaussian splatting has emerged as a powerful tool for high-fidelity reconstruction of dynamic scenes. However, existing methods primarily rely on implicit motion representations, s…
Bounding Distributional Shifts in World Modeling through Novelty Detection
Eric Jing, Abdeslam Boularias
Recent work on visual world models shows significant promise in latent state dynamics obtained from pre-trained image backbones. However, most of the current approaches are sensiti…