80 citations · 251 across the 17 of their papers we have counts for
8 papers · 1 filter
Benchmarking Open-Ended Multi-Agent Coordination in Language Agents
Kale-ab Abebe Tessera, Andras Szecsenyi, Cameron Barker +7
As language models are increasingly deployed as autonomous agents, they must coordinate with others over long horizons in open-ended interactive tasks. Yet existing evaluations rar…
LLM-First Search: Self-Guided Exploration of the Solution Space
Nathan Herr, Tim Rocktäschel, Roberta Raileanu
Large Language Models (LLMs) have demonstrated remarkable improvements in reasoning and planning through increased test-time compute, often by framing problem-solving as a search p…
TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
Jonathan Cook, Tim Rocktäschel, Jakob Foerster +2
Given the widespread adoption and usage of Large Language Models (LLMs), it is crucial to have flexible and interpretable evaluations of their instruction-following ability. Prefer…
Learning Reasoning Strategies in End-to-End Differentiable Proving
Pasquale Minervini, Sebastian Riedel, Pontus Stenetorp +2
Attempts to render deep learning models interpretable, data-efficient, and robust have seen some success through hybridisation with rule-based systems, for example, in Neural Theor…
WordCraft: An Environment for Benchmarking Commonsense Agents
Minqi Jiang, Jelena Luketina, Nantas Nardelli +4
The ability to quickly solve a wide range of real-world tasks requires a commonsense understanding of the world. Yet, how to best extract such knowledge from natural language corpo…
Generating Interactive Worlds with Text
Angela Fan, Jack Urbanek, Pratik Ringshia +8
Procedurally generating cohesive and interesting game environments is challenging and time-consuming. In order for the relationships between the game elements to be natural, common…