20 papers · 1 filter
cVLA: Towards Efficient Camera-Space VLAs
Max Argus, Jelena Bratulic, Houman Masnavi +4
Vision-Language-Action (VLA) models offer a compelling framework for tackling complex robotic manipulation tasks, but they are often expensive to train. In this paper, we propose a…
CoVAR: Co-generation of Video and Action for Robotic Manipulation via Multi-Modal Diffusion
Liudi Yang, Yang Bai, George Eskandar +5
We present a method to generate video-action pairs that follow text instructions, starting from an initial image observation and the robot's joint states. Our approach automaticall…
Efficient Learning of Object Placement with Intra-Category Transfer
Adrian Röfer, Russell Buchanan, Max Argus +2
Efficient learning from demonstration for long-horizon tasks remains an open challenge in robotics. While significant effort has been directed toward learning trajectories, a recen…
Enhancing LLM-based Autonomous Driving with Modular Traffic Light and Sign Recognition
Fabian Schmidt, Noushiq Mohammed Kayilan Abdul Nazar, Markus Enzweiler +1
Large Language Models (LLMs) are increasingly used for decision-making and planning in autonomous driving, showing promising reasoning capabilities and potential to generalize acro…
OpenLex3D: A Tiered Evaluation Benchmark for Open-Vocabulary 3D Scene Representations
Christina Kassab, Sacha Morin, Martin Büchner +5
3D scene understanding has been transformed by open-vocabulary language models that enable interaction via natural language. However, at present the evaluation of these representat…
Preference-Based Long-Horizon Robotic Stacking with Multimodal Large Language Models
Wanming Yu, Adrian Röfer, Abhinav Valada +1
Pretrained large language models (LLMs) can work as high-level robotic planners by reasoning over abstract task descriptions and natural language instructions, etc. However, they h…