Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
MathDuels: A Self-Play Benchmark That Grows
Zhiqiu Xu, Shibo Jin, Shreya Arya +1
As frontier language models attain near-ceiling performance on static mathematical benchmarks, existing evaluations are increasingly unable to differentiate model capabilities, lar…
cs.CL2025
Once Upon an Input: Reasoning via Per-Instance Program Synthesis
Adam Stein, Neelay Velingker, Mayur Naik +1
Large language models (LLMs) excel at zero-shot inference but continue to struggle with complex, multi-step reasoning. Recent methods that augment LLMs with intermediate reasoning…
cs.CL2024
Towards Compositionality in Concept Learning
Adam Stein, Aaditya Naik, Yinjun Wu +2
Concept-based interpretability methods offer a lens into the internals of foundation models by decomposing their embeddings into high-level concepts. These concept representations…