activity
20242026
collaborators

6 papers

cs.RO2026

Hybrid Training for Vision-Language-Action Models

Pietro Mazzaglia, Cansu Sancaktar, Markus Peschl +1

Using Large Language Models to produce intermediate thoughts, a.k.a. Chain-of-thought (CoT), before providing an answer has been a successful recipe for solving complex language ta…

cs.LG2026

Multi-scale Predictive Representations for Goal-conditioned Reinforcement Learning

Valliappan Chidambaram Adaikkappan, David Meger, Sai Rajeswar +1

This paper investigates robust representation learning in offline goal-conditioned reinforcement learning (GCRL). Particularly in sparse reward scenarios, learning representations…

cs.AI2026

LUMINA: Long-horizon Understanding for Multi-turn Interactive Agents

Amin Rakhsha, Thomas Hehn, Pietro Mazzaglia +3

Large language models can perform well on many isolated tasks, yet they continue to struggle on multi-turn, long-horizon agentic problems that require skills such as planning, stat…

cs.RO2025

From Code to Action: Hierarchical Learning of Diffusion-VLM Policies

Markus Peschl, Pietro Mazzaglia, Daniel Dijkman

Imitation learning for robotic manipulation often suffers from limited generalization and data scarcity, especially in complex, long-horizon tasks. In this work, we introduce a hie…

cs.RO2025

Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models

Rokas Bendikas, Daniel Dijkman, Markus Peschl +2

Vision-Language-Action (VLA) models offer a pivotal approach to learning robotic manipulation at scale by repurposing large pre-trained Vision-Language-Models (VLM) to output robot…

cs.RO2024

Redundancy-aware Action Spaces for Robot Learning

Pietro Mazzaglia, Nicholas Backshall, Xiao Ma +1

Joint space and task space control are the two dominant action modes for controlling robot arms within the robot learning literature. Actions in joint space provide precise control…