4 papers
GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization
Xiaosong Jia, Bowen Yang, Zuhao Ge +17
Vision-Language-Action (VLA) models aim for general robot learning by aligning action as a modality within powerful Vision-Language Models (VLMs). Existing VLAs rely on end-to-end…
AssemMate: Graph-Based LLM for Robotic Assembly Assistance
Qi Zheng, Chaoran Zhang, Zijian Liang +5
Large Language Model (LLM)-based robotic assembly assistance has gained significant research attention. It requires the injection of domain-specific knowledge to guide the assembly…
DSM: Constructing a Diverse Semantic Map for 3D Visual Grounding
Qinghongbing Xie, Zijian Liang, Fuhao Li +1
Effective scene representation is critical for the visual grounding ability of representations, yet existing methods for 3D Visual Grounding are often constrained. They either only…
Visual Agentic Reinforcement Fine-Tuning
Ziyu Liu, Yuhang Zang, Yushan Zou +6
A key trend in Large Reasoning Models (e.g., OpenAI's o3) is the native agentic ability to use external tools such as web browsers for searching and writing/executing code for imag…