3 papers
cs.AI2026
Benchmarking Prompt Optimization of Large Language Models With Chess
Timothée Lesort, Alejandra López de Aberasturi Gómez, Tristan Karch +4
Evaluating large language models becomes increasingly challenging as their capabilities advance: benchmarks can saturate, public test sets risk contamination, and assessing harder…
cs.CL2026
Future Querying: Can LLMs Serve as Implicit Medical World Models?
Siri Willems, James Butterworth, Lore Goetschalckx +4
Traditional clinical prediction models rely on task-specific pipelines and curated, structured data, which scale poorly and underutilize unstructured text. To address this, we intr…
cs.AI2025
Surfer-H Meets Holo1: Cost-Efficient Web Agent Powered by Open Weights
Mathieu Andreux, Breno Baldas Skuk, Hamza Benchekroun +41
We present Surfer-H, a cost-efficient web agent that integrates Vision-Language Models (VLM) to perform user-defined tasks on the web. We pair it with Holo1, a new open-weight coll…