25 papers
RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?
Zhenxuan Fan, Bo Zhang, Yutong Lin +9
Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task compl…
EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents
Wei Wang, Wenqiao Zhang, Yutong Lin +14
Vision-language-action (VLA) models map visual observations and language instructions directly to robot actions, but long-horizon tasks require more than action prediction. An agen…
Holtercare-Bench: A Multimodal Benchmark for Evaluating Long-Term Dynamic ECG Analysis
Yihan Xie, Hanwen Cui, Runze Ye +8
While multimodal large language models (MLLMs) excel in medical applications, most of them favor static images or short-term signals. In the critical field of dynamic electrocardio…
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning
Haoyu Zheng, Yun Zhu, Qing Wang +1
Recent agentic reinforcement learning methods use hindsight to complement sparse outcome rewards. However, a completed rollout can yield many such signals, leaving their appropriat…
E-MRL: Cross-view Aligned Evidence-driven Multimodal Reinforcement Learning for Reliable 3D Tumor Analysis
Sijing Li, Zhongwei Qiu, Zhuoya Wang +6
While Vision-Language Models (VLMs) show great promise in volumetric medical report generation, they frequently suffer from visual hallucinations and a lack of grounding in 3D CT d…
SCOPE: Evolving Symbolic World for Planning in Open-Ended Environments
Yundaichuan Zhan, Minghe Gao, Zhongqi Yue +7
Recent works have explored integrating Vision-Language Models (VLMs) with classical planners that rely on symbolic representations of planning problems to generate long-horizon pla…