5 papers
RePO-VLA: Recovery-Driven Policy Optimization for Vision-Language-Action Models
Weijia Liufu, Xiaoyu Guo, Ruiyi Chen +16
Vision-Language-Action (VLA) models remain brittle in long-horizon, contact-rich manipulation because success-only imitation provides little supervision for execution drift, while…
ArchGPT: Understanding the World's Architectures with Large Multimodal Models
Yuze Wang, Luo Yang, Junyi Wang +1
Architecture embodies aesthetic, cultural, and historical values, standing as a tangible testament to human civilization. Researchers have long leveraged virtual reality (VR), mixe…
Taking Language Embedded 3D Gaussian Splatting into the Wild
Yuze Wang, Yue Qi
Recent advances in leveraging large-scale Internet photo collections for 3D reconstruction have enabled immersive virtual exploration of landmarks and historic sites worldwide. How…
Seg-Wild: Interactive Segmentation based on 3D Gaussian Splatting for Unconstrained Image Collections
Yongtang Bao, Chengjie Tang, Yuze Wang +1
Reconstructing and segmenting scenes from unconstrained photo collections obtained from the Internet is a novel but challenging task. Unconstrained photo collections are easier to…
DeOcc-1-to-3: 3D De-Occlusion from a Single Image via Self-Supervised Multi-View Diffusion
Yansong Qu, Shaohui Dai, Xinyang Li +4
Reconstructing 3D objects from a single image remains challenging, especially under real-world occlusions. While recent diffusion-based view synthesis models can generate consisten…