1 citations · 1 across the 2 of their papers we have counts for
3 papers
Search-Time Data Contamination
Ziwen Han, Meher Mankikar, Julian Michael +1
Data contamination refers to the leakage of evaluation data into model training data, resulting in overfitting to supposedly held-out test sets and compromising test validity. We i…
FORTRESS: Frontier Risk Evaluation for National Security and Public Safety
Christina Q. Knight, Kaustubh Deshpande, Ved Sirdeshmukh +4
The rapid advancement of large language models (LLMs) introduces dual-use capabilities that could both threaten and bolster national security and public safety (NSPS). Models imple…
A Red Teaming Roadmap Towards System-Level Safety
Zifan Wang, Christina Q. Knight, Jeremy Kritz +2
Large Language Model (LLM) safeguards, which implement request refusals, have become a widely adopted mitigation strategy against misuse. At the intersection of adversarial machine…