4 papers
Who Judges the Judges? A Chinese Safety QA Benchmark for Evaluating LLM Responses and Safety Judges
Rui Yang, Shuang Huang, Junhua Liu +7
Safety benchmarks for large language models often assess the risk of a user query, although the outcome of question answering depends on whether the response violates a policy. Thi…
Active Testing of Large Language Models via Approximate Neyman Allocation
Zeli Liu, Jiancheng Zhang, Cong Liu +1
Large language models (LLMs) require reliable evaluation from pre-training to test-time scaling, making evaluation a recurring rather than one-off cost. As model scales grow and ta…
PPM: Automated Generation of Diverse Programming Problems for Benchmarking Code Generation Models
Simin Chen, Xiaoning Feng, Xiaohong Han +2
In recent times, a plethora of Large Code Generation Models (LCGMs) have been proposed, showcasing significant potential in assisting developers with complex programming tasks. Ben…
Uncertainty Awareness of Large Language Models Under Code Distribution Shifts: A Benchmark Study
Yufei Li, Simin Chen, Yanghong Guo +3
Large Language Models (LLMs) have been widely employed in programming language analysis to enhance human productivity. Yet, their reliability can be compromised by various code dis…