activity
20242026
collaborators

5 papers

cs.AI2026

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents

Kale-ab Abebe Tessera, Andras Szecsenyi, Cameron Barker +7

As language models are increasingly deployed as autonomous agents, they must coordinate with others over long horizons in open-ended interactive tasks. Yet existing evaluations rar…

cs.LG2026

Déjà Q: Open-Ended Evolution of Diverse, Learnable and Verifiable Problems

Willem Röpke, Samuel Coward, Andrei Lupu +3

Recent advances in reasoning models have yielded impressive results in mathematics and coding. However, most approaches rely on static datasets, which have been suggested to encour…

cs.AI2025

LLM-First Search: Self-Guided Exploration of the Solution Space

Nathan Herr, Tim Rocktäschel, Roberta Raileanu

Large Language Models (LLMs) have demonstrated remarkable improvements in reasoning and planning through increased test-time compute, often by framing problem-solving as a search p…

cs.LG2025

Preference-Based Alignment of Discrete Diffusion Models

Umberto Borso, Davide Paglieri, Jude Wells +1

Diffusion models have achieved state-of-the-art performance across multiple domains, with recent advancements extending their applicability to discrete data. However, aligning disc…

cs.AI2024

TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation

Jonathan Cook, Tim Rocktäschel, Jakob Foerster +2

Given the widespread adoption and usage of Large Language Models (LLMs), it is crucial to have flexible and interpretable evaluations of their instruction-following ability. Prefer…