activity
20242026
collaborators
Showing 2025Show all

20 papers · 1 filter

cs.RO2025

cVLA: Towards Efficient Camera-Space VLAs

Max Argus, Jelena Bratulic, Houman Masnavi +4

Vision-Language-Action (VLA) models offer a compelling framework for tackling complex robotic manipulation tasks, but they are often expensive to train. In this paper, we propose a…

cs.CV2025

CoVAR: Co-generation of Video and Action for Robotic Manipulation via Multi-Modal Diffusion

Liudi Yang, Yang Bai, George Eskandar +5

We present a method to generate video-action pairs that follow text instructions, starting from an initial image observation and the robot's joint states. Our approach automaticall…

cs.RO2025

Efficient Learning of Object Placement with Intra-Category Transfer

Adrian Röfer, Russell Buchanan, Max Argus +2

Efficient learning from demonstration for long-horizon tasks remains an open challenge in robotics. While significant effort has been directed toward learning trajectories, a recen…

cs.CV2025

Enhancing LLM-based Autonomous Driving with Modular Traffic Light and Sign Recognition

Fabian Schmidt, Noushiq Mohammed Kayilan Abdul Nazar, Markus Enzweiler +1

Large Language Models (LLMs) are increasingly used for decision-making and planning in autonomous driving, showing promising reasoning capabilities and potential to generalize acro…

cs.CV2025

OpenLex3D: A Tiered Evaluation Benchmark for Open-Vocabulary 3D Scene Representations

Christina Kassab, Sacha Morin, Martin Büchner +5

3D scene understanding has been transformed by open-vocabulary language models that enable interaction via natural language. However, at present the evaluation of these representat…

cs.RO2025

Preference-Based Long-Horizon Robotic Stacking with Multimodal Large Language Models

Wanming Yu, Adrian Röfer, Abhinav Valada +1

Pretrained large language models (LLMs) can work as high-level robotic planners by reasoning over abstract task descriptions and natural language instructions, etc. However, they h…