2 papers
cs.RO2026
AsyncMDE: Real-Time Monocular Depth Estimation via Asynchronous Spatial Memory
Lianjie Ma, Yuquan Li, Bingzheng Jiang +3
Foundation-model-based monocular depth estimation offers a viable alternative to active sensors for robot perception, yet its computational cost often prohibits deployment on edge…
cs.RO2026
DepthCache: Depth-Guided Training-Free Visual Token Merging for Vision-Language-Action Model Inference
Yuquan Li, Lianjie Ma, Han Ding +1
Vision-Language-Action (VLA) models enable generalist robotic manipulation but suffer from high inference latency. This bottleneck stems from the massive number of visual tokens pr…