Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Learning to Reason Efficiently with A* Post-Training
Andreas Opedal, Francesco Ignazio Re, Abulhair Saparov +3
Many applications of large language models (LLMs) require deductive reasoning, yet models frequently produce incorrect or redundant inference steps. We frame natural language infer…
cs.AI2026
Test of Time: Rethinking Temporal Signal of Benchmark Contamination
Terry Jingchen Zhang, Gopal Dev, Ning Wang +8
Post-cutoff performance decay of LLMs has been widely interpreted as a temporal signal for benchmark contamination, where public information released before the training cutoff may…