Publications (14)
Neighborhood Mixup Experience Replay: Local Convex Interpolation for Improved Sample Efficiency in Continuous Control Tasks
Ryan Sander, Wilko Schwarting, Tim Seyde +3
Experience replay plays a crucial role in improving the sample efficiency of deep reinforcement learning agents. Recent advances in experience replay propose using Mixup (Zhang et…
Zero-Overhead Introspection for Adaptive Test-Time Compute
Rohin Manvi, Joey Hong, Tim Seyde +3
Large language models excel at reasoning but lack key aspects of introspection, including anticipating their own success and the computation required to achieve it. Humans use real…
In-Place Tokenizer Expansion for Pre-trained LLMs
Jimmy T. H. Smith, Tarek Dakhran, Alberto Cabrera +7
The paper proposes an in‑place tokenizer expansion method that continues a pre‑trained model’s BPE merges on multilingual data, reuses existing token embeddings, and initializes ne…
Towards Cooperative Flight Control Using Visual-Attention
Lianhao Yin, Makram Chahine, Tsun-Hsuan Wang +5
The cooperation of a human pilot with an autonomous agent during flight control realizes parallel autonomy. We propose an air-guardian system that facilitates cooperation between a…
Faster Algorithms for Growing Collision-Free Convex Polytopes in Robot Configuration Space
Peter Werner, Thomas Cohn, Rebecca H. Jiang +4
We propose two novel algorithms for constructing convex collision-free polytopes in robot configuration space. Finding these polytopes enables the application of stronger motion-pl…
Solving Continuous Control via Q-learning
Tim Seyde, Peter Werner, Wilko Schwarting +4
While there has been substantial success for solving continuous control with actor-critic methods, simpler critic-only methods such as Q-learning find limited application in the as…
Is Bang-Bang Control All You Need? Solving Continuous Control with Bernoulli Policies
Tim Seyde, Igor Gilitschenski, Wilko Schwarting +4
Reinforcement learning (RL) for continuous control typically employs distributions whose support covers the entire action space. In this work, we investigate the colloquially known…
Locomotion Planning through a Hybrid Bayesian Trajectory Optimization
Tim Seyde, Jan Carius, Ruben Grandia +2
Locomotion planning for legged systems requires reasoning about suitable contact schedules. The contact sequence and timings constitute a hybrid dynamical system and prescribe a su…
Deep Latent Competition: Learning to Race Using Visual Control Policies in Latent Space
Wilko Schwarting, Tim Seyde, Igor Gilitschenski +4
Learning competitive behaviors in multi-agent settings such as racing requires long-term reasoning about potential adversarial interactions. This paper presents Deep Latent Competi…
LFM2 Technical Report
Alexander Amini, Anna Banaszak, Harold Benoit +30
We present LFM2, a family of Liquid Foundation Models designed for efficient on-device deployment and strong task capabilities. Using hardware-in-the-loop architecture search under…
Interpreting Neural Policies with Disentangled Tree Representations
Tsun-Hsuan Wang, Wei Xiao, Tim Seyde +2
The advancement of robots, particularly those functioning in complex human-centric environments, relies on control solutions that are driven by machine learning. Understanding how…
Learning to Plan Optimistically: Uncertainty-Guided Deep Exploration via Latent Model Ensembles
Tim Seyde, Wilko Schwarting, Sertac Karaman +1
Learning complex robot behaviors through interaction requires structured exploration. Planning should target interactions with the potential to optimize long-term performance, whil…
Multi-Agent Robotic Control with Onboard Vision-Language Models
Kajetan RachwaÅ, Maciej Majek, BartÅomiej Boczek +6
Vision Language Models (VLMs) and Vision Language Action (VLA) models have shown promise in robotic control. Yet, they face significant challenges regarding explainability, general…
Growing Q-Networks: Solving Continuous Control Tasks with Adaptive Control Resolution
Tim Seyde, Peter Werner, Wilko Schwarting +2
Recent reinforcement learning approaches have shown surprisingly strong capabilities of bang-bang policies for solving continuous control benchmarks. The underlying coarse action s…