Showing cs.ROShow all
2 papers · 1 filter
cs.RO2026
VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis
Xiaolei Lang, Yang Wang, Yukun Zhou +10
Recent advances in robot foundation models trained on large-scale human teleoperation data have enabled robots to perform increasingly complex real-world tasks. However, scaling th…
cs.RO2024
Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks
Pranav Guruprasad, Harshvardhan Sikka, Jaewoo Song +2
Vision-language-action (VLA) models represent a promising direction for developing general-purpose robotic systems, demonstrating the ability to combine visual understanding, langu…