4 papers
MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent
Yuxia Fu, Zhizhen Zhang, Yuqi Zhang +3
Recent Vision-Language-Action (VLA) models reformulate vision-language models by tuning them with millions of robotic demonstrations. While they perform well when fine-tuned for a…
The RoboSense Challenge: Sense Anything, Navigate Anywhere, Adapt Across Platforms
Lingdong Kong, Shaoyuan Xie, Zeying Gong +135
Autonomous systems are increasingly deployed in open and dynamic environments -- from city streets to aerial and indoor spaces -- where perception models must remain reliable under…
Box-QAymo: Box-Referring VQA Dataset for Autonomous Driving
Djamahl Etchegaray, Yuxia Fu, Zi Huang +1
Interpretable communication is essential for safe and trustworthy autonomous driving, yet current vision-language models (VLMs) often operate under idealized assumptions and strugg…
SCORE: Soft Label Compression-Centric Dataset Condensation via Coding Rate Optimization
Bowen Yuan, Yuxia Fu, Zijian Wang +2
Dataset Condensation (DC) aims to obtain a condensed dataset that allows models trained on the condensed dataset to achieve performance comparable to those trained on the full data…