From the 2 of 15 linked papers with an AI index.
15 papers
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models
Yufeng Ji, Wenhao Tang, Haoyi Niu +3
The paper introduces Action QFormer, a query-based interface that reorganizes multimodal information into action-focused representations to improve vision-language-action models, e…
Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents
Yixian Zhang, Huanming Zhang, Feng Gao +13
The paper introduces Harness VLA, a memory-augmented framework that combines a frozen vision‑language‑action model with a small set of analytic manipulation primitives to improve r…
LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies
Jialei Chen, Kai Wang, Kang Chen +9
Vision-Language-Action models (VLAs) leverage large-scale vision-language pretraining for semantic robot control, but often lack explicit foresight into how robot actions change th…
Mastering Multi-Drone Volleyball through Hierarchical Co-Self-Play Reinforcement Learning
Ruize Zhang, Sirui Xiang, Zelai Xu +6
In this paper, we tackle the problem of learning to play 3v3 multi-drone volleyball, a new embodied competitive task that requires both high-level strategic coordination and low-le…
VolleyBots: A Testbed for Multi-Drone Volleyball Game Combining Motion Control and Strategic Play
Zelai Xu, Ruize Zhang, Chao Yu +9
Robot sports, characterized by well-defined objectives, explicit rules, and dynamic interactions, present ideal scenarios for demonstrating embodied intelligence. In this paper, we…
RLinf-USER: A Unified and Extensible System for Real-World Online Policy Learning in Embodied AI
Hongzhi Zang, Shu'ang Yu, Hao Lin +14
Online policy learning directly in the physical world is a promising yet challenging direction for embodied intelligence. Unlike simulation, real-world systems cannot be arbitraril…