6 citations · 7 across the 7 of their papers we have counts for
3 papers · 1 filter
Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models
Phongsakon Mark Konrad, Toygar Tanyel, Serkan Ayvaz
Safety evaluations often assume that behavior observed during testing reflects behavior in ordinary use, but fine-tuning can break this assumption. A checkpoint can appear fixed un…
The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime
Phongsakon Mark Konrad, Tim Lukas Adam, Ane Cathrine Holst Merrild +4
AI deployment in sensitive domains such as health care, credit, employment, and criminal justice is often treated as unsafe to authorize until model internals can be explained. Thi…
Acceptance Cards:A Four-Diagnostic Standard for Safe Fine-Tuning Defense Claims
Phongsakon Mark Konrad, Toygar Tanyel, Serkan Ayvaz
Safe fine-tuning defenses are often endorsed on the basis of a held-out gap reduction, but the same reduction can come from sampling noise, subject artifacts, capability loss, or a…