2 papers
cs.AI2026
Alignment Inertia: Auditing the Durability of Training Data Influence Through Policy Override Resistance
Renata Barreto, Markelle Roesti, Mohammad Tahaei
Platform operators increasingly rely on system prompts and fine-tuning to govern model behavior, yet it remains unclear how reliably these interventions override behavior inherited…
cs.HC2026
Who Does What in AI Auditing? Designing Human-AI Collaboration for Auditing Generative AI
Eunkyu Park, Markelle Roesti, Wesley Hanwen Deng +5
AI auditing increasingly incorporates AI agents to expand the scale and breadth of audit coverage, yet little is known about how auditing work should be divided without displacing…