works on

From the 1 of 5 linked papers with an AI index.

activity
20242026
collaborators

5 papers

math.CT2026

Left properness of Moore flows

Philippe Gaucher

We introduce the notion of a reparametrization category with cuts. For every such reparametrization category , we prove the tensor lemma, namely that the tensor product…

cs.LG2026

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

Xin Qiu, Yulu Gan, Conor F. Hayes +6

The paper shows that evolution strategies can successfully fine‑tune billion‑parameter large language models without backpropagation, outperforming reinforcement learning in stabil…

cs.LG2026

Overcoming Forgetting in LLM Fine-Tuning with Evolution Strategies

Kajetan Schweighofer, Conor F. Hayes, Roberto Dailey +2

Evolution Strategies (ES) has recently emerged as a competitive alternative to reinforcement learning (RL) for large language model (LLM) fine-tuning, offering advantages through s…

cs.AI2025

Solving a Million-Step LLM Task with Zero Errors

Elliot Meyerson, Giuseppe Paolo, Roberto Dailey +6

LLMs have achieved remarkable breakthroughs in reasoning, insights, and tool use, but chaining these abilities into extended processes at the scale of those routinely executed by h…

cs.LG2024

Multi-objective Reinforcement Learning: A Tool for Pluralistic Alignment

Peter Vamplew, Conor F Hayes, Cameron Foale +2

Reinforcement learning (RL) is a valuable tool for the creation of AI systems. However it may be problematic to adequately align RL based on scalar rewards if there are multiple co…