7 papers · 1 filter
Instance-Optimal Estimation with Multiple LLM Judges on a Budget
Junghyun Lee, Sanghwa Kim, Yassir Jedra +2
Evaluating large language models increasingly relies on LLM-as-a-judge protocols, but such evaluations remain costly: different judges have different prices and reliabilities, and…
Curvature-Guided LoRA: Matching Full Fine-Tuning in Function Space
Frédéric Zheng, Alexandre Proutière
Parameter-efficient fine-tuning methods such as LoRA enable efficient adaptation of large pretrained models, but often lag behind full fine-tuning in both convergence speed and fin…
Shift Before You Learn: Enabling Low-Rank Representations in Reinforcement Learning
Bastien Dubail, Stefan Stojanovic, Alexandre Proutière
Low-rank structure is a common implicit assumption in many modern reinforcement learning (RL) algorithms. For instance, reward-free and goal-conditioned RL methods often presume th…
Model-free Low-Rank Reinforcement Learning via Leveraged Entry-wise Matrix Estimation
Stefan Stojanovic, Yassir Jedra, Alexandre Proutiere
We consider the problem of learning an -optimal policy in controlled dynamical systems with low-rank latent structure. For this problem, we present LoRa-PI (Low-Rank P…
Conformal Predictions under Markovian Data
Frédéric Zheng, Alexandre Proutiere
We study the split Conformal Prediction method when applied to Markovian data. We quantify the gap in terms of coverage induced by the correlations in the data (compared to exchang…
Low-Rank Bandits via Tight Two-to-Infinity Singular Subspace Recovery
Yassir Jedra, William Réveillard, Stefan Stojanovic +1
We study contextual bandits with low-rank structure where, in each round, if the (context, arm) pair is selected, the learner observes a noisy sample of th…