3 papers
cs.AI2026
Benchmarking Prompt Optimization of Large Language Models With Chess
Timothée Lesort, Alejandra López de Aberasturi Gómez, Tristan Karch +4
Evaluating large language models becomes increasingly challenging as their capabilities advance: benchmarks can saturate, public test sets risk contamination, and assessing harder…
cs.SE2025
LLM Agents for Interactive Exploration of Historical Cadastre Data: Framework and Application to Venice
Tristan Karch, Jakhongir Saydaliev, Isabella Di Lenardo +1
Cadastral data reveal key information about the historical organization of cities but are often non-standardized due to diverse formats and human annotations, complicating large-sc…
cs.CL2025
Is This Collection Worth My LLM's Time? Automatically Measuring Information Potential in Text Corpora
Tristan Karch, Luca Engel, Philippe Schwaller +1
As large language models (LLMs) converge towards similar capabilities, the key to advancing their performance lies in identifying and incorporating valuable new information sources…