3 papers
cs.RO2025
VLA-RAIL: A Real-Time Asynchronous Inference Linker for VLA Models and Robots
Yongsheng Zhao, Lei Zhao, Baoping Cheng +3
Vision-Language-Action (VLA) models have achieved remarkable breakthroughs in robotics, with the action chunk playing a dominant role in these advances. Given the real-time and con…
cs.GR2025
MoDA: Multi-modal Diffusion Architecture for Talking Head Generation
Xinyang Li, Gen Li, Zhihui Lin +6
Talking head generation with arbitrary identities and speech audio remains a crucial problem in the realm of the virtual metaverse. Recently, diffusion models have become a popular…
cs.CV2024
Monocular Visual Place Recognition in LiDAR Maps via Cross-Modal State Space Model and Multi-View Matching
Gongxin Yao, Xinyang Li, Luowei Fu +1
Achieving monocular camera localization within pre-built LiDAR maps can bypass the simultaneous mapping process of visual SLAM systems, potentially reducing the computational overh…