works on

From the 1 of 7 linked papers with an AI index.

collaborators

7 papers

cs.AI2026

CIPHER: A Decoupled Exploration-Selection Framework for Test-Time Scaling of Data Science Agents

Maxime Heuillet, Sharadind Peddiraju

CIPHER is an AI-powered data science agent that improves reliability by generating multiple candidate initial states and selecting the best ones for parallel execution at test time…

cs.LG2026

Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts

Maxime Heuillet, Yufei Cui, Boxing Chen +2

Advanced reasoning in LLMs on challenging domains like mathematical reasoning can be tackled using verifiable rewards based reinforced fine-tuning (ReFT). In standard ReFT framewor…

cs.LG2026

Neural Active Learning Meets the Partial Monitoring Framework

Maxime Heuillet, Ola Ahmad, Audrey Durand

We focus on the online-based active learning (OAL) setting where an agent operates over a stream of observations and trades-off between the costly acquisition of information (label…

cs.LG2026

Randomized Confidence Bounds for Stochastic Partial Monitoring

Maxime Heuillet, Ola Ahmad, Audrey Durand

The partial monitoring (PM) framework provides a theoretical formulation of sequential learning problems with incomplete feedback. On each round, a learning agent plays an action w…

cs.LG2026

Robust Fine-Tuning from Non-Robust Pretrained Models: Mitigating Suboptimal Transfer With Epsilon-Scheduling

Jonas Ngnawé, Maxime Heuillet, Sabyasachi Sahoo +5

Fine-tuning pretrained models is a standard and effective workflow in modern machine learning. However, robust fine-tuning (RFT), which aims to simultaneously achieve adaptation to…

cs.IR2025

LLM-as-a-Judge: Toward World Models for Slate Recommendation Systems

Baptiste Bonin, Maxime Heuillet, Audrey Durand

Modeling user preferences across domains remains a key challenge in slate recommendation (i.e. recommending an ordered sequence of items) research. We investigate how Large Languag…