5 papers
Towards Precise Intent-Aligned VLA Aerial Navigation via Expert-Guided GRPO
Tianyang Chen, Wenjun Li, Xin zhou +2
Vision-Language-Action (VLA) models offer a promising end-to-end paradigm for unmanned aerial vehicles (UAVs) to accomplish complex tasks specified by fine-grained instructions. Ho…
Vision-Language-Action in Robotics: A Survey of Datasets, Benchmarks, and Data Engines
Ziyao Wang, Bingying Wang, Hanrong Zhang +7
Despite remarkable progress in Vision--Language--Action (VLA) models, a central bottleneck remains underexamined: the data infrastructure that underlies embodied learning. In this…
Precise Aggressive Aerial Maneuvers with Sensorimotor Policies
Tianyue Wu, Guangtong Xu, Zihan Wang +6
Precise aggressive maneuvers with lightweight onboard sensors remains a key bottleneck in fully exploiting the maneuverability of drones. Such maneuvers are critical for expanding…
Towards Closed-Loop Embodied Empathy Evolution: Probing LLM-Centric Lifelong Empathic Motion Generation in Unseen Scenarios
Jiawen Wang, Jingjing Wang Tianyang Chen, Min Zhang +1
In the literature, existing human-centric emotional motion generation methods primarily focus on boosting performance within a single scale-fixed dataset, largely neglecting the fl…
GroundSight: Augmenting Vision-Language Models with Grounding Information and De-hallucination
Xinxi Chen, Tianyang Chen, Lijia Hong
We propose a method to improve Visual Question Answering (VQA) with Retrieval-Augmented Generation (RAG) by introducing text-grounded object localization. Rather than retrieving in…