2 papers
cs.CL2026
MathDuels: A Self-Play Benchmark That Grows
Zhiqiu Xu, Shibo Jin, Shreya Arya +1
As frontier language models attain near-ceiling performance on static mathematical benchmarks, existing evaluations are increasingly unable to differentiate model capabilities, lar…
cs.CV2025
B-cosification: Transforming Deep Neural Networks to be Inherently Interpretable
Shreyash Arya, Sukrut Rao, Moritz Böhle +1
B-cos Networks have been shown to be effective for obtaining highly human interpretable explanations of model decisions by architecturally enforcing stronger alignment between inpu…