1 paper
Ralf Römer, Maximilian Seeliger, Saida Liu +5
Vision-language-action models (VLAs) combine vision-language backbones with expressive generative action heads trained via flow matching on large-scale robotic datasets. Despite th…