3 citations · 4 across the 5 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving
Xiyuan Zhou, Zhuoqi Li, Xinlei Wang +6
Data contamination undermines the reliable evaluation of large language models (LLMs) on mathematical problem solving. While rewriting-based evaluation mitigates memorization, exis…
cs.CL2025
Multi-Objective Large Language Model Unlearning
Zibin Pan, Shuwen Zhang, Yuesheng Zheng +3
Machine unlearning in the domain of large language models (LLMs) has attracted great attention recently, which aims to effectively eliminate undesirable behaviors from LLMs without…