Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks
Junlin Liu, Shengnan An, Shuang Zhou +10
Contemporary large language models (LLMs) have demonstrated remarkable reasoning capabilities, particularly in specialized domains like mathematics and physics. However, their abil…
cs.CL2025
AMO-Bench: Large Language Models Still Struggle in High School Math Competitions
Shengnan An, Xunliang Cai, Xuezhi Cao +8
We present AMO-Bench, an Advanced Mathematical reasoning benchmark with Olympiad level or even higher difficulty, comprising 50 human-crafted problems. Existing benchmarks have wid…
cs.CL2024
DetoxBench: Benchmarking Large Language Models for Multitask Fraud & Abuse Detection
Joymallya Chakraborty, Wei Xia, Anirban Majumder +3
Large language models (LLMs) have demonstrated remarkable capabilities in natural language processing tasks. However, their practical application in high-stake domains, such as fra…