From the 1 of 10 linked papers with an AI index.
10 papers
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
Hengyi Xie, Chenfei Yao, Xianjin Wu +7
TurboVLA is a vision-language-action model that directly maps visual observations and language instructions to robot actions, achieving real-time performance (32 Hz) on an RTX 4090…
AsyncMDE: Real-Time Monocular Depth Estimation via Asynchronous Spatial Memory
Lianjie Ma, Yuquan Li, Bingzheng Jiang +3
Foundation-model-based monocular depth estimation offers a viable alternative to active sensors for robot perception, yet its computational cost often prohibits deployment on edge…
MoRI: Mixture of RL and IL Experts for Long-Horizon Manipulation Tasks
Yaohang Xu, Lianjie Ma, Gewei Zuo +3
Reinforcement Learning (RL) and Imitation Learning (IL) are the standard frameworks for policy acquisition in manipulation. While IL offers efficient policy derivation, it suffers…
DepthCache: Depth-Guided Training-Free Visual Token Merging for Vision-Language-Action Model Inference
Yuquan Li, Lianjie Ma, Han Ding +1
Vision-Language-Action (VLA) models enable generalist robotic manipulation but suffer from high inference latency. This bottleneck stems from the massive number of visual tokens pr…
Omnidirectional Humanoid Locomotion on Stairs via Unsafe Stepping Penalty and Sparse LiDAR Elevation Mapping
Yuzhi Jiang, Yujun Liang, Junhao Li +2
Humanoid robots, characterized by numerous degrees of freedom and a high center of gravity, are inherently unstable. Safe omnidirectional locomotion on stairs requires both omnidir…
A Three-Level Whole-Body Disturbance Rejection Control Framework for Dynamic Motions in Legged Robots
Bolin Li, Gewei Zuo, Zhixiang Wang +3
This paper presents a control framework designed to enhance the stability and robustness of legged robots in the presence of uncertainties, including model uncertainties, external…