24 papers
GOPI: Generation-Oriented 3D Pose Inference for Furniture Insertion from Single-View RGB-D Indoor Scenes
Ruifeng Zhai, Renjie Liu, Guangrun Wang +1
We study the problem of inserting new furniture into indoor scene images. Under masked single-view 2D image-plane conditioning, however, the physical scale of the inserted furnitur…
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute
Hongyu Chen, Liang Lin, Guangrun Wang
The paper proposes Self‑Verifying Refinement (SVR), a reinforcement‑learning framework that lets language models decide when to stop refining answers by using their own correctness…
Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation
Zijian Song, Qichang Li, Sihan Qin +4
The scarcity of large-scale robotic data has motivated the repurposing of foundation models from other modalities for policy learning. In this work, we introduce PhysGen (Learning…
Learning Spatial-Temporal Coherent Correlations for Speech-Preserving Facial Expression Manipulation
Tianshui Chen, Jianman Lin, Zhijing Yang +3
Speech-preserving facial expression manipulation (SPFEM) aims to modify facial emotions while meticulously maintaining the mouth animation associated with spoken content. Current w…
Stable Language Guidance for Vision-Language-Action Models
Zhihao Zhan, Yuhao Chen, Jiaying Zhou +5
Vision-Language-Action (VLA) models have demonstrated impressive capabilities in generalized robotic control; however, they remain notoriously brittle to linguistic perturbations.…
Robotic Manipulation is Vision-to-Geometry Mapping: Vision-Geometry Backbones over Language and Video Models
Zijian Song, Qichang Li, Jiawei Zhou +4
At its core, robotic manipulation is a problem of vision-to-geometry mapping (). Physical actions are fundamentally defined by geometric properties like 3D posi…