4 papers
GOPI: Generation-Oriented 3D Pose Inference for Furniture Insertion from Single-View RGB-D Indoor Scenes
Ruifeng Zhai, Renjie Liu, Guangrun Wang +1
We study the problem of inserting new furniture into indoor scene images. Under masked single-view 2D image-plane conditioning, however, the physical scale of the inserted furnitur…
VLA Models Are More Generalizable Than You Think: Revisiting Physical and Spatial Modeling
Weiqi Li, Quande Zhang, Ruifeng Zhai +2
Vision-language-action (VLA) models achieve strong in-distribution performance but degrade sharply under novel camera viewpoints and visual perturbations. We show that this brittle…
TAG: Target-Agnostic Guidance for Stable Object-Centric Inference in Vision-Language-Action Models
Jiaying Zhou, Zhihao Zhan, Ruifeng Zhai +5
Vision--Language--Action (VLA) policies have shown strong progress in mapping language instructions and visual observations to robotic actions, yet their reliability degrades in cl…
E0: Enhancing Generalization and Fine-Grained Control in VLA Models via Tweedie Discrete Diffusion
Zhihao Zhan, Jiaying Zhou, Likui Zhang +10
Vision-Language-Action (VLA) models offer a unified framework for robotic manipulation by integrating visual perception, language understanding, and control generation. However, ex…