Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning
Nearchos Potamitis, Vansh Ramani, Har Ashish Arora +3
Benchmark scores for LLM reasoning systems are reported as single numbers, yet the same model, strategy, and task can produce meaningfully different answers and costs across repeat…
cs.AI2026
Agentic Proving for Program Verification
Alessandro Sosso, Akhil Arora, Bas Spitters
Agentic systems have recently emerged as state-of-the-art approaches for automated theorem proving in formal mathematics. To assess how far these capabilities extend to program ver…