Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Does Reasoning Improve Psychological Depth in Large Language Models? It Depends on Who's Judging
Ruichen Zheng, Yihe Wang, Fabrice Y Harel-Canada +4
LLM-as-a-Judge evaluators are increasingly used to score open-ended generation, yet a judge's correlation with human ratings on its development set may not guarantee valid measurem…
cs.LG2026
TaoBench: Do Automated Theorem Prover LLMs Generalize Beyond MathLib?
Alexander K Taylor, Junyi Zhang, Ethan Ji +10
Automated theorem proving (ATP) benchmarks largely consist of problems formalized in MathLib, so current ATP training and evaluation are heavily biased toward MathLib's definitiona…