From the 1 of 7 linked papers with an AI index.
7 papers
CIPHER: A Decoupled Exploration-Selection Framework for Test-Time Scaling of Data Science Agents
Maxime Heuillet, Sharadind Peddiraju
CIPHER is an AI-powered data science agent that improves reliability by generating multiple candidate initial states and selecting the best ones for parallel execution at test time…
Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts
Maxime Heuillet, Yufei Cui, Boxing Chen +2
Advanced reasoning in LLMs on challenging domains like mathematical reasoning can be tackled using verifiable rewards based reinforced fine-tuning (ReFT). In standard ReFT framewor…
Neural Active Learning Meets the Partial Monitoring Framework
Maxime Heuillet, Ola Ahmad, Audrey Durand
We focus on the online-based active learning (OAL) setting where an agent operates over a stream of observations and trades-off between the costly acquisition of information (label…
Randomized Confidence Bounds for Stochastic Partial Monitoring
Maxime Heuillet, Ola Ahmad, Audrey Durand
The partial monitoring (PM) framework provides a theoretical formulation of sequential learning problems with incomplete feedback. On each round, a learning agent plays an action w…
Robust Fine-Tuning from Non-Robust Pretrained Models: Mitigating Suboptimal Transfer With Epsilon-Scheduling
Jonas Ngnawé, Maxime Heuillet, Sabyasachi Sahoo +5
Fine-tuning pretrained models is a standard and effective workflow in modern machine learning. However, robust fine-tuning (RFT), which aims to simultaneously achieve adaptation to…
LLM-as-a-Judge: Toward World Models for Slate Recommendation Systems
Baptiste Bonin, Maxime Heuillet, Audrey Durand
Modeling user preferences across domains remains a key challenge in slate recommendation (i.e. recommending an ordered sequence of items) research. We investigate how Large Languag…