3 papers
cs.AI2025
Search-Time Data Contamination
Ziwen Han, Meher Mankikar, Julian Michael +1
Data contamination refers to the leakage of evaluation data into model training data, resulting in overfitting to supposedly held-out test sets and compromising test validity. We i…
cs.CL2025
MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs
Alexander R. Fabbri, Diego Mares, Jorge Flores +5
Although recent Large Language Models (LLMs) have shown rapid improvement on reasoning benchmarks in English, the evaluation of such LLMs' multilingual reasoning capability across…
cs.CY2025
FORTRESS: Frontier Risk Evaluation for National Security and Public Safety
Christina Q. Knight, Kaustubh Deshpande, Ved Sirdeshmukh +4
The rapid advancement of large language models (LLMs) introduces dual-use capabilities that could both threaten and bolster national security and public safety (NSPS). Models imple…