2 papers
cs.RO2026
EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents
Wei Wang, Wenqiao Zhang, Yutong Lin +14
Vision-language-action (VLA) models map visual observations and language instructions directly to robot actions, but long-horizon tasks require more than action prediction. An agen…
cs.LG2026
Holtercare-Bench: A Multimodal Benchmark for Evaluating Long-Term Dynamic ECG Analysis
Yihan Xie, Hanwen Cui, Runze Ye +8
While multimodal large language models (MLLMs) excel in medical applications, most of them favor static images or short-term signals. In the critical field of dynamic electrocardio…