activity
20152026
most citedSafe Policy Search for Lifelong Reinforcement Learning with Sublinear Regret

21 citations · 78 across the 29 of their papers we have counts for

collaborators
Showing cs.AIShow all

9 papers · 1 filter

cs.AI2026

Risk-Controlled Lean-as-Judge for Natural-Language Mathematical Reasoning

Pauline Bourigault, Xiaotong Ji, Matthieu Zimmer +2

Lean is increasingly used to judge natural-language mathematical answers, but its signal is partial: many answers never formalize, and a failed proof may reflect an ill-typed state…

cs.AI2026

A Brain-like Synergistic Core in LLMs Drives Behaviour and Learning

Pedro Urbina-Rodriguez, Zafeirios Fountas, Fernando E. Rosas +5

The independent evolution of intelligence in biological and artificial systems offers a unique opportunity to identify its fundamental computational principles. Here we show that l…

cs.AI2025

Model-Based and Sample-Efficient AI-Assisted Math Discovery in Sphere Packing

Rasul Tutunov, Alexandre Maraval, Antoine Grosnit +3

Sphere packing, Hilbert's eighteenth problem, asks for the densest arrangement of congruent spheres in n-dimensional Euclidean space. Although relevant to areas such as cryptograph…

cs.AI2025

Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving

Matthieu Zimmer, Xiaotong Ji, Rasul Tutunov +3

Reasoning remains a challenging task for large language models (LLMs), especially within the logically constrained environment of automated theorem proving (ATP), due to sparse rew…

cs.AI2024

Human-inspired Episodic Memory for Infinite Context LLMs

Zafeirios Fountas, Martin A Benfeghoul, Adnan Oomerjee +4

Large language models (LLMs) have shown remarkable capabilities, but still struggle with processing extensive contexts, limiting their ability to maintain coherence and accuracy ov…

cs.AI2023

Pangu-Agent: A Fine-Tunable Generalist Agent with Structured Reasoning

Filippos Christianos, Georgios Papoudakis, Matthieu Zimmer +13

A key method for creating Artificial Intelligence (AI) agents is Reinforcement Learning (RL). However, constructing a standalone RL policy that maps perception to action directly e…