works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.LG2026

Models Can Model, But Can't Bind: Structured Grounding in Text-to-Optimization

Zhiqi Gao, Albert Ge, Alexander Berenbeim +2

The paper investigates why text‑to‑optimization models struggle to correctly ground problem data, introduces a benchmark (Text2Opt‑Bench) to study this, and proposes a binding‑focu…

cs.LG2026

Counterexample Guided Learning in the Large using Reasoning Agents

Hongyi Liu, Frederic Sala, Thomas Reps +1

LLMs and LLM agents should improve when given feedback, but identifying when they are able to do so is difficult: feedback is heterogeneous, domain-specific, and difficult to contr…

cs.LG2026

Fine-Tuning Small Reasoning Models for Quantum Field Theory

Nathaniel S. Woodward, Zhiqi Gao, Yurii Kvasiuk +3

Despite the growing application of Large Language Models (LLMs) to theoretical physics, there is little academic exploration into how domain-specific physics reasoning ability deve…

cs.LG2025

Test-time Scaling Techniques in Theoretical Physics -- A Comparison of Methods on the TPBench Dataset

Zhiqi Gao, Tianyi Li, Yurii Kvasiuk +5

Large language models (LLMs) have shown strong capabilities in complex reasoning, and test-time scaling techniques can enhance their performance with comparably low cost. Many of t…

cs.LG2025

Theoretical Physics Benchmark (TPBench) -- a Dataset and Study of AI Reasoning Capabilities in Theoretical Physics

Daniel J. H. Chung, Zhiqi Gao, Yurii Kvasiuk +5

We introduce a benchmark to evaluate the capability of AI to solve problems in theoretical physics, focusing on high-energy theory and cosmology. The first iteration of our benchma…