4 papers
Privacy-Preserving AI Verification via Minimal Information Disclosure
Sleem Abdelghafar, Gabriel Kulp
AI verification crosses a trust boundary: a verifier must learn enough to establish an authorized claim, yet the same evidence can reveal sensitive details about the model, workloa…
Expert Selections In MoE Models Reveal (Almost) As Much As Text
Amir Nuriyev, Gabriel Kulp
We present a text-reconstruction attack on mixture-of-experts (MoE) language models that recovers tokens from expert selections alone. In MoE models, each token is routed to a subs…
Verifying International Agreements on AI: Six Layers of Verification for Rules on Large-Scale AI Development and Deployment
Mauricio Baker, Gabriel Kulp, Oliver Marks +2
The risks of frontier AI may require international cooperation, which in turn may require verification: checking that all parties follow agreed-on rules. For instance, states might…
Hardware-Enabled Mechanisms for Verifying Responsible AI Development
Aidan O'Gara, Gabriel Kulp, Will Hodgkins +7
Advancements in AI capabilities, driven in large part by scaling up computing resources used for AI training, have created opportunities to address major global challenges but also…