Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Test-time Scaling Techniques in Theoretical Physics -- A Comparison of Methods on the TPBench Dataset
Zhiqi Gao, Tianyi Li, Yurii Kvasiuk +5
Large language models (LLMs) have shown strong capabilities in complex reasoning, and test-time scaling techniques can enhance their performance with comparably low cost. Many of t…
cs.LG2025
Theoretical Physics Benchmark (TPBench) -- a Dataset and Study of AI Reasoning Capabilities in Theoretical Physics
Daniel J. H. Chung, Zhiqi Gao, Yurii Kvasiuk +5
We introduce a benchmark to evaluate the capability of AI to solve problems in theoretical physics, focusing on high-energy theory and cosmology. The first iteration of our benchma…