2 papers
cs.CL2026
MathDuels: A Self-Play Benchmark That Grows
Zhiqiu Xu, Shibo Jin, Shreya Arya +1
As frontier language models attain near-ceiling performance on static mathematical benchmarks, existing evaluations are increasingly unable to differentiate model capabilities, lar…
cs.LG2026
On Improving Neurosymbolic Learning by Exploiting the Representation Space
Aaditya Naik, Efthymia Tsamoura, Shibo Jin +2
We study the problem of learning neural classifiers in a neurosymbolic setting where the hidden gold labels of input instances must satisfy a logical formula. Learning in this sett…