3 papers
cs.LG2026
Sharding Prevents LLM Oversight Failures and Adversarial Exploitation
Victor Akinwande, J. Zico Kolter, Aran Nayebi
Giving an LLM judge more compute does not necessarily make it check more requirements. When one call must return many verdicts, some decisions become weakly grounded in the evidenc…
cs.CV2024
HyperCLIP: Adapting Vision-Language models with Hypernetworks
Victor Akinwande, Mohammad Sadegh Norouzzadeh, Devin Willmott +3
Self-supervised vision-language models trained with contrastive objectives form the basis of current state-of-the-art methods in AI vision tasks. The success of these models is a d…
cs.CL2024
Introducing v0.5 of the AI Safety Benchmark from MLCommons
Bertie Vidgen, Adarsh Agrawal, Ahmed M. Ahmed +97
This paper introduces v0.5 of the AI Safety Benchmark, which has been created by the MLCommons AI Safety Working Group. The AI Safety Benchmark has been designed to assess the safe…