2 papers
cs.CL2024
MARIO Eval: Evaluate Your Math LLM with your Math LLM--A mathematical dataset evaluation toolkit
Boning Zhang, Chengxi Li, Kai Fan
Large language models (LLMs) have been explored in a variety of reasoning tasks including solving of mathematical problems. Each math dataset typically includes its own specially d…
eess.AS2022
I4U System Description for NIST SRE'20 CTS Challenge
Kong Aik Lee, Tomi Kinnunen, Daniele Colibro +23
This manuscript describes the I4U submission to the 2020 NIST Speaker Recognition Evaluation (SRE'20) Conversational Telephone Speech (CTS) Challenge. The I4U's submission was resu…