2 papers
cs.CL2026
Beyond Repeated Sampling: Learning Search Policies for LLM Reasoning
Ismail Labiad, Matthieu Kowalski, Marc Schoenauer +2
Large language models increasingly tackle hard reasoning problems by spending more test-time compute, yet the dominant strategy remains naive repeated sampling: draw many independe…
cs.LG2025
Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training
Ismail Labiad, Mathurin Videau, Matthieu Kowalski +4
Gradient-based optimization is the workhorse of deep learning, offering efficient and scalable training via backpropagation. However, exposing gradients during training can leak se…