5 papers
Benchmarking Open-Ended Multi-Agent Coordination in Language Agents
Kale-ab Abebe Tessera, Andras Szecsenyi, Cameron Barker +7
As language models are increasingly deployed as autonomous agents, they must coordinate with others over long horizons in open-ended interactive tasks. Yet existing evaluations rar…
Déjà Q: Open-Ended Evolution of Diverse, Learnable and Verifiable Problems
Willem Röpke, Samuel Coward, Andrei Lupu +3
Recent advances in reasoning models have yielded impressive results in mathematics and coding. However, most approaches rely on static datasets, which have been suggested to encour…
LLM-First Search: Self-Guided Exploration of the Solution Space
Nathan Herr, Tim Rocktäschel, Roberta Raileanu
Large Language Models (LLMs) have demonstrated remarkable improvements in reasoning and planning through increased test-time compute, often by framing problem-solving as a search p…
Preference-Based Alignment of Discrete Diffusion Models
Umberto Borso, Davide Paglieri, Jude Wells +1
Diffusion models have achieved state-of-the-art performance across multiple domains, with recent advancements extending their applicability to discrete data. However, aligning disc…
TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
Jonathan Cook, Tim Rocktäschel, Jakob Foerster +2
Given the widespread adoption and usage of Large Language Models (LLMs), it is crucial to have flexible and interpretable evaluations of their instruction-following ability. Prefer…