7 papers
Co-VLA: Coordination-Aware Structured Action Modeling for Dual-Arm Vision-Language-Action Systems
Yandong Wang, Jiaqian Yu, Xiongfeng Peng +8
Vision-language-action (VLA) models show strong capabilities in single and dual-arm robotic manipulation. Prior works show coordinated bimanual behaviors can emerge from end-to-end…
DAM-VLA: A Dynamic Action Model-Based Vision-Language-Action Framework for Robot Manipulation
Xiongfeng Peng, Jiaqian Yu, Dingzhe Li +8
In dynamic environments such as warehouses, hospitals, and homes, robots must seamlessly transition between gross motion and precise manipulations to complete complex tasks. Howeve…
Diffusion Transformer meets Multi-level Wavelet Spectrum for Single Image Super-Resolution
Peng Du, Hui Li, Han Xu +5
Discrete Wavelet Transform (DWT) has been widely explored to enhance the performance of image superresolution (SR). Despite some DWT-based methods improving SR by capturing fine-gr…
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines
Lu Xu, Jiaqian Yu, Xiongfeng Peng +7
To meet the growing demand for smarter, faster, and more efficient embodied AI solutions, we introduce a novel Mixture-of-Expert (MoE) method that significantly boosts reasoning an…
3D Occupancy Prediction with Low-Resolution Queries via Prototype-aware View Transformation
Gyeongrok Oh, Sungjune Kim, Heeju Ko +7
The resolution of voxel queries significantly influences the quality of view transformation in camera-based 3D occupancy prediction. However, computational constraints and the prac…
Unveiling the Hidden: Online Vectorized HD Map Construction with Clip-Level Token Interaction and Propagation
Nayeon Kim, Hongje Seong, Daehyun Ji +1
Predicting and constructing road geometric information (e.g., lane lines, road markers) is a crucial task for safe autonomous driving, while such static map elements can be repeate…