Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
StatEval: A Comprehensive Benchmark for Large Language Models in Statistics
Yuchen Lu, Run Yang, Yichen Zhang +6
Despite rapid advances in large language models (LLMs), statistical reasoning remains underrepresented in existing LLM benchmarks, which often do not reflect the layered, proof-dri…
cs.CL2025
R1-RE: Cross-Domain Relation Extraction with RLVR
Runpeng Dai, Tong Zheng, Run Yang +2
Relation extraction (RE) is a core task in natural language processing. Traditional approaches typically frame RE as a supervised learning problem, directly mapping context to labe…
cs.CL2025
Enhancing Clinical Multiple-Choice Questions Benchmarks with Knowledge Graph Guided Distractor Generation
Running Yang, Wenlong Deng, Minghui Chen +2
Clinical tasks such as diagnosis and treatment require strong decision-making abilities, highlighting the importance of rigorous evaluation benchmarks to assess the reliability of…