activity
20202026
most citedSMARTS: Scalable Multi-Agent Reinforcement Learning Training School for Autonomous Driving

103 citations · 118 across the 15 of their papers we have counts for

collaborators
Showing cs.LGShow all

9 papers · 1 filter

cs.LG2025

Subjective Depth and Timescale Transformers: Learning Where and When to Compute

Frederico Wieser, Martin Benfeghoul, Haitham Bou Ammar +2

The rigid, uniform allocation of computation in standard Transformer (TF) architectures can limit their efficiency and scalability, particularly for large-scale models and long seq…

cs.LG2025

Untangling Component Imbalance in Hybrid Linear Attention Conversion Methods

Martin Benfeghoul, Teresa Delgado, Adnan Oomerjee +3

Transformers' quadratic computational complexity limits their scalability despite remarkable performance. While linear attention reduces this to linear complexity, pre-training suc…

cs.LG2025

Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective

Matthieu Zimmer, Xiaotong Ji, Tu Nguyen +1

We introduce a novel approach to large language model (LLM) distillation by formulating it as a constrained reinforcement learning problem. While recent work has begun exploring th…

cs.LG2025

On Almost Surely Safe Alignment of Large Language Models at Inference-Time

Xiaotong Ji, Shyam Sundhar Ramesh, Matthieu Zimmer +3

We introduce a novel inference-time alignment approach for LLMs that aims to generate safe responses almost surely, i.e., with probability approaching one. Our approach models the…

cs.LG2024

Efficient Reinforcement Learning with Large Language Model Priors

Xue Yan, Yan Song, Xidong Feng +4

In sequential decision-making (SDM) tasks, methods like reinforcement learning (RL) and heuristic search have made notable advances in specific cases. However, they often require e…

cs.LG2023

End-to-End Meta-Bayesian Optimisation with Transformer Neural Processes

Alexandre Maraval, Matthieu Zimmer, Antoine Grosnit +1

Meta-Bayesian optimisation (meta-BO) aims to improve the sample efficiency of Bayesian optimisation by leveraging data from related tasks. While previous methods successfully meta-…