7 papers
LA4VLA: Learning to Act without Seeing via Language-Action Pretraining
Tao Lin, Yuxin Du, Yiran Mao +13
Vision-Language-Action (VLA) models are commonly pretrained on robot demonstrations by jointly mapping visual observations and language instructions to actions. However, dense visu…
Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model
Tao Lin, Yuxin Du, Jiting Liu +14
Vision-Language-Action models have emerged as a promising paradigm for robotic manipulation by unifying perception, language grounding, and action generation. However, they often s…
Focusable Monocular Depth Estimation
Yuxin Du, Tao Lin, Zile Zhong +7
Monocular depth foundation models generalize well across scenes, yet they are typically optimized with uniform pixel-wise objectives that do not distinguish user-specified or task-…
Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
Tao Lin, Yilei Zhong, Yuxin Du +11
Vision-Language-Action (VLA) models have emerged as a powerful framework that unifies perception, language, and control, enabling robots to perform diverse tasks through multimodal…
Multi-modal and Metadata Capture Model for Micro Video Popularity Prediction
Jiacheng Lu, Mingyuan Xiao, Weijian Wang +3
As short videos have become the primary form of content consumption across various industries, accurately predicting their popularity has become key to enhancing user engagement an…
SegVol: Universal and Interactive Volumetric Medical Image Segmentation
Yuxin Du, Fan Bai, Tiejun Huang +1
Precise image segmentation provides clinical study with instructive information. Despite the remarkable progress achieved in medical image segmentation, there is still an absence o…