5 papers
ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models
Wencheng Ye, Tianshi Wang, Lei Zhu +3
Recent Vision-Language-Action (VLA) models have shown impressive flexibility and generalization, yet their deployment in robotic manipulation remains limited by heavy computational…
Non-Markovian Long-Horizon Robot Manipulation via Keyframe Chaining
Yipeng Chen, Wentao Tan, Lei Zhu +4
Existing Vision-Language-Action (VLA) models often struggle to generalize to long-horizon tasks due to their heavy reliance on immediate observations. While recent studies incorpor…
MOTIF: Learning Action Motifs for Few-shot Cross-Embodiment Transfer
Heng Zhi, Wentao Tan, Lei Zhu +4
While vision-language-action (VLA) models have advanced generalist robotic learning, cross-embodiment transfer remains challenging due to kinematic heterogeneity and the high cost…
Dynamic Embedding of Hierarchical Visual Features for Efficient Vision-Language Fine-Tuning
Xinyu Wei, Guoli Yang, Jialu Zhou +4
Large Vision-Language Models (LVLMs) commonly follow a paradigm that projects visual features and then concatenates them with text tokens to form a unified sequence input for Large…
MapAgent: Trajectory-Constructed Memory-Augmented Planning for Mobile Task Automation
Yi Kong, Dianxi Shi, Guoli Yang +4
The recent advancement of autonomous agents powered by Large Language Models (LLMs) has demonstrated significant potential for automating tasks on mobile devices through graphical…