2 papers
cs.CL2026
evalci: A Python Library for Statistically Rigorous Comparison of Language Model Evaluations
Shreyas K Chandrahas
The dominant practice in language model evaluation is to report a single accuracy number per model and declare the higher one better, without testing whether the gap could plausibl…
cs.LG2026
Constrained Tabular Diffusion for Finance
Michael Cardei, Jose M Munoz, Oscar Barrera +2
Generative models in finance face the dual challenge of producing realistic data while satisfying strict regulatory and economic objectives, a requirement that standard tabular dif…