32 citations · 77 across the 35 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs?
Juliusz Ziomek, William Bankes, Lorenz Wolf +3
We introduce LLM-Wikirace, a benchmark for evaluating planning, reasoning, and world knowledge in large language models (LLMs). In LLM-Wikirace, models must efficiently navigate Wi…
cs.AI2025
Imagined Autocurricula
Ahmet H. Güzel, Matthew Thomas Jackson, Jarek Luca Liesen +4
Training agents to act in embodied environments typically requires vast training data or access to accurate simulation, neither of which exists for many cases in the real world. In…