From the 1 of 5 linked papers with an AI index.
5 papers
Models Can Model, But Can't Bind: Structured Grounding in Text-to-Optimization
Zhiqi Gao, Albert Ge, Alexander Berenbeim +2
The paper investigates why text‑to‑optimization models struggle to correctly ground problem data, introduces a benchmark (Text2Opt‑Bench) to study this, and proposes a binding‑focu…
Counterexample Guided Learning in the Large using Reasoning Agents
Hongyi Liu, Frederic Sala, Thomas Reps +1
LLMs and LLM agents should improve when given feedback, but identifying when they are able to do so is difficult: feedback is heterogeneous, domain-specific, and difficult to contr…
Fine-Tuning Small Reasoning Models for Quantum Field Theory
Nathaniel S. Woodward, Zhiqi Gao, Yurii Kvasiuk +3
Despite the growing application of Large Language Models (LLMs) to theoretical physics, there is little academic exploration into how domain-specific physics reasoning ability deve…
Test-time Scaling Techniques in Theoretical Physics -- A Comparison of Methods on the TPBench Dataset
Zhiqi Gao, Tianyi Li, Yurii Kvasiuk +5
Large language models (LLMs) have shown strong capabilities in complex reasoning, and test-time scaling techniques can enhance their performance with comparably low cost. Many of t…
Theoretical Physics Benchmark (TPBench) -- a Dataset and Study of AI Reasoning Capabilities in Theoretical Physics
Daniel J. H. Chung, Zhiqi Gao, Yurii Kvasiuk +5
We introduce a benchmark to evaluate the capability of AI to solve problems in theoretical physics, focusing on high-energy theory and cosmology. The first iteration of our benchma…