From the 1 of 10 linked papers with an AI index.
10 papers
RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination
Haotian Liang, Mingkang Chen, Yufei Huang +27
The paper introduces RxBrain, a foundation model that jointly reasons over language and visual inputs to create embodied plans, using a multimodal Mixture-of-Transformers architect…
4DSloMo: 4D Reconstruction for High Speed Scene with Asynchronous Capture
Yutian Chen, Shi Guo, Tianshuo Yang +4
Reconstructing fast-dynamic scenes from multi-view videos is crucial for high-speed motion analysis and realistic 4D reconstruction. However, the majority of 4D capture systems are…
Rein3D: Reinforced 3D Indoor Scene Generation with Panoramic Video Diffusion Models
Dehui Wang, Rong Wei, Yue Shi +9
The growing demand for Embodied AI and VR applications has highlighted the need for synthesizing high-quality 3D indoor scenes from sparse inputs. However, existing approaches stru…
Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
Zhixuan Liang, Yizhuo Li, Tianshuo Yang +9
Vision-Language-Action (VLA) models adapt large vision-language backbones to map images and instructions into robot actions. However, prevailing VLAs either generate actions autore…
HiVLA: A Visual-Grounded-Centric Hierarchical Embodied Manipulation System
Tianshuo Yang, Guanyu Chen, Yutian Chen +8
While end-to-end Vision-Language-Action (VLA) models offer a promising paradigm for robotic manipulation, fine-tuning them on narrow control data often compromises the profound rea…
AnyRecon: Arbitrary-View 3D Reconstruction with Video Diffusion Model
Yutian Chen, Shi Guo, Renbiao Jin +7
Sparse-view 3D reconstruction is essential for modeling scenes from casual captures, but remain challenging for non-generative reconstruction. Existing diffusion-based approaches m…