#financial NLP
try —
2 papers match
cs.CL2026
APEX-Accounting
Julien Benchek, Austin Bennett, Jasmin Kern +8
The paper presents APEX-Accounting, a benchmark for evaluating how well advanced language models can perform real accounting tasks such as reconciliation, expense accrual, transact…
#accounting automation#benchmark#large language models#evaluation metrics
cs.AI2026
Measurement Risk in Supervised Financial NLP: Rubric and Metric Sensitivity on JF-ICR
Sidi Chang, Peiying Zhu, Yuxiao Chen +1
The paper investigates how the wording of evaluation rubrics and the choice of metrics affect the reliability of supervised financial NLP benchmarks, using a Japanese implicit‑comm…
#financial nlp#benchmark evaluation#rubric sensitivity#metric selection