5 papers
StatEval: A Comprehensive Benchmark for Large Language Models in Statistics
Yuchen Lu, Run Yang, Yichen Zhang +6
Despite rapid advances in large language models (LLMs), statistical reasoning remains underrepresented in existing LLM benchmarks, which often do not reflect the layered, proof-dri…
R1-RE: Cross-Domain Relation Extraction with RLVR
Runpeng Dai, Tong Zheng, Run Yang +2
Relation extraction (RE) is a core task in natural language processing. Traditional approaches typically frame RE as a supervised learning problem, directly mapping context to labe…
Breach in the Shield: Unveiling the Vulnerabilities of Large Language Models
Runpeng Dai, Run Yang, Fan Zhou +1
Large Language Models (LLMs) and Vision-Language Models (VLMs) have achieved impressive performance across a wide range of tasks, yet they remain vulnerable to carefully crafted pe…
Enhancing Clinical Multiple-Choice Questions Benchmarks with Knowledge Graph Guided Distractor Generation
Running Yang, Wenlong Deng, Minghui Chen +2
Clinical tasks such as diagnosis and treatment require strong decision-making abilities, highlighting the importance of rigorous evaluation benchmarks to assess the reliability of…
Spatio-temporal Prediction of Fine-Grained Origin-Destination Matrices with Applications in Ridesharing
Run Yang, Runpeng Dai, Siran Gao +3
Accurate spatial-temporal prediction of network-based travelers' requests is crucial for the effective policy design of ridesharing platforms. Having knowledge of the total demand…