activity
20232026
collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2026

Instance-Optimal Estimation with Multiple LLM Judges on a Budget

Junghyun Lee, Sanghwa Kim, Yassir Jedra +2

Evaluating large language models increasingly relies on LLM-as-a-judge protocols, but such evaluations remain costly: different judges have different prices and reliabilities, and…

cs.LG2026

Curvature-Guided LoRA: Matching Full Fine-Tuning in Function Space

Frédéric Zheng, Alexandre Proutière

Parameter-efficient fine-tuning methods such as LoRA enable efficient adaptation of large pretrained models, but often lag behind full fine-tuning in both convergence speed and fin…

cs.LG2025

Shift Before You Learn: Enabling Low-Rank Representations in Reinforcement Learning

Bastien Dubail, Stefan Stojanovic, Alexandre Proutière

Low-rank structure is a common implicit assumption in many modern reinforcement learning (RL) algorithms. For instance, reward-free and goal-conditioned RL methods often presume th…

cs.LG2024

Model-free Low-Rank Reinforcement Learning via Leveraged Entry-wise Matrix Estimation

Stefan Stojanovic, Yassir Jedra, Alexandre Proutiere

We consider the problem of learning an -optimal policy in controlled dynamical systems with low-rank latent structure. For this problem, we present LoRa-PI (Low-Rank P…

cs.LG2024

Conformal Predictions under Markovian Data

Frédéric Zheng, Alexandre Proutiere

We study the split Conformal Prediction method when applied to Markovian data. We quantify the gap in terms of coverage induced by the correlations in the data (compared to exchang…

cs.LG2024

Low-Rank Bandits via Tight Two-to-Infinity Singular Subspace Recovery

Yassir Jedra, William Réveillard, Stefan Stojanovic +1

We study contextual bandits with low-rank structure where, in each round, if the (context, arm) pair is selected, the learner observes a noisy sample of th…