2 papers
cs.AI2026
How Inference Compute Shapes Frontier LLM Evaluation
Jessica McFadyen, Ole Jorgensen, Harry Coppock +2
The paper studies how the amount of compute allocated during inference (e.g., token budget, repeated attempts) affects the performance of frontier large language models on challeng…
cs.CL2025
MedCaseReasoning: Evaluating and learning diagnostic reasoning from clinical case reports
Kevin Wu, Eric Wu, Rahul Thapa +7
Doctors and patients alike increasingly use Large Language Models (LLMs) to diagnose clinical cases. However, unlike domains such as math or coding, where correctness can be object…