7 papers
SignVLA: Real-Time Sign Language-Guided Robotic Manipulation via Attention LSTM and Vision-Language-Action Models
Ningwei Bai, Xinyu Tan, Harry Gardner +6
Vision-Language-Action (VLA) models enable robots to execute manipulation tasks from natural-language instructions grounded in visual observations. However, existing VLA interfaces…
Elastic Queries Reinforcement Learning: Self-Aware Policy Execution for VLA Models
Ge Wang, Xinyu Tan, Xiang Li +11
Vision-language-action (VLA) models are powerful action generators for robot manipulation, but they are typically executed with fixed inference and replanning schedules. This rigid…
MUSE: Agentic 3D Scene Authoring via Memory-Grounded Incremental Requirement Satisfaction
Ruijie Xu, Xinnan Zhu, Jiayu Ying +3
Text-driven 3D scene generation is a promising technique for digital content creation, embodied AI simulation, and interactive design, yet practical workflows often require refinin…
GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents
Xiongbin Wu, Zhihao Luo, Shanzhe Lei +7
Recently, vision-language model (VLM) agents have shown promising progress in open-world tasks, where successful task completion often requires multiple turns of visual perception…
Formal Skill: Programmable Runtime Skills for Efficient and Accurate LLM Agents
Xi Zhang, Meijun Gao, Yuntian Zhao +6
Large Language Model (LLM) agents increasingly act inside real workspaces, where tools and skills determine whether model reasoning becomes reliable action. Existing skills remain…
SignVLA: A Gloss-Free Vision-Language-Action Framework for Real-Time Sign Language-Guided Robotic Manipulation
Xinyu Tan, Ningwei Bai, Harry Gardener +6
We present, to our knowledge, the first sign language-driven Vision-Language-Action (VLA) framework for intuitive and inclusive human-robot interaction. Unlike conventional approac…