9 papers
Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation
Junjin Xiao, Dongyang Li, Yandan Yang +9
This paper tackles spatial perception and manipulation challenges in Vision-Language-Action (VLA) models. To address depth ambiguity from monocular input, we leverage a pre-trained…
World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training
Junjin Xiao, Yandan Yang, Xinyuan Chang +5
Vision-Language-Action (VLA) models trained via imitation learning suffer from significant performance degradation in data-scarce scenarios due to their reliance on large-scale dem…
Generative Texture Filtering
Rongjia Zheng, Shangwei Huang, Lei Zhu +2
We present a generative method for texture filtering, which exhibits surprisingly good performance and generalizability. Our core idea is to empower texture filtering by taking ful…
SGS-Intrinsic: Semantic-Invariant Gaussian Splatting for Sparse-View Indoor Inverse Rendering
Jiahao Niu, Rongjia Zheng, Wenju Xu +2
We present SGS-Intrinsic, an indoor inverse rendering framework that works well for sparse-view images. Unlike existing 3D Gaussian Splatting (3DGS) based methods that focus on obj…
You Only Erase Once: Erasing Anything without Bringing Unexpected Content
Yixing Zhu, Qing Zhang, Wenju Xu +1
We present YOEO, an approach for object erasure. Unlike recent diffusion-based methods which struggle to erase target objects without generating unexpected content within the maske…
Learning Explicit Continuous Motion Representation for Dynamic Gaussian Splatting from Monocular Videos
Xuankai Zhang, Junjin Xiao, Shangwei Huang +2
We present an approach for high-quality dynamic Gaussian Splatting from monocular videos. To this end, we in this work go one step further beyond previous methods to explicitly mod…