4 papers
Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model
Tao Lin, Yuxin Du, Jiting Liu +14
Vision-Language-Action models have emerged as a promising paradigm for robotic manipulation by unifying perception, language grounding, and action generation. However, they often s…
Focusable Monocular Depth Estimation
Yuxin Du, Tao Lin, Zile Zhong +7
Monocular depth foundation models generalize well across scenes, yet they are typically optimized with uniform pixel-wise objectives that do not distinguish user-specified or task-…
CAST: Collapse-Aware multi-Scale Topology Fusion for Multimodal Coreset Selection
Boran Zhao, Hetian Liu, Zhenxian Hu +3
The training of large multimodal models fundamentally relies on massive image-text datasets, which inevitably incur prohibitive computational overhead. Dataset selection offers a p…
ECHO: Continuous Hierarchical Memory for Vision-Language-Action Models
Yanbin Hu, Jin Cui, Jiayi Lu +6
Memory capacity is a critical factor determining the performance of Vision-Language-Action (VLA) models in long-horizon manipulation tasks. Existing memory-augmented architectures…