Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
How Long Is a Piece of String? A Brief Empirical Analysis of Tokenizers
Jonathan Roberts, Kai Han, Samuel Albanie
Frontier LLMs are increasingly utilised across academia, society and industry. A commonly used unit for comparing models, their inputs and outputs, and estimating inference pricing…
cs.CL2025
GAMEBoT: Transparent Assessment of LLM Reasoning in Games
Wenye Lin, Jonathan Roberts, Yunhan Yang +3
Large Language Models (LLMs) are increasingly deployed in real-world applications that demand complex reasoning. To track progress, robust benchmarks are required to evaluate their…
cs.CL2025
Needle Threading: Can LLMs Follow Threads through Near-Million-Scale Haystacks?
Jonathan Roberts, Kai Han, Samuel Albanie
As the context limits of Large Language Models (LLMs) increase, the range of possible applications and downstream functions broadens. In many real-world tasks, decisions depend on…