Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
EXECUTE: A Multilingual Benchmark for LLM Token Understanding
Lukas Edman, Helmut Schmid, Alexander Fraser
The CUTE benchmark showed that LLMs struggle with character understanding in English. We extend it to more languages with diverse scripts and writing systems, introducing EXECUTE.…
cs.CL2024
CUTE: Measuring LLMs' Understanding of Their Tokens
Lukas Edman, Helmut Schmid, Alexander Fraser
Large Language Models (LLMs) show remarkable performance on a wide variety of tasks. Most LLMs split text into multi-character tokens and process them as atomic units without direc…