2 papers
cs.CR2026
Aggressive Compression Enables LLM Weight Theft
Davis Brown, Juan-Pablo Rivera, Dan Hendrycks +1
As frontier AIs become more powerful and costly to develop, adversaries have increasing incentives to steal model weights by mounting exfiltration attacks. In this work, we conside…
cs.LG2024
Towards Measuring Goal-Directedness in AI Systems
Dylan Xu, Juan-Pablo Rivera
Recent advances in deep learning have brought attention to the possibility of creating advanced, general AI systems that outperform humans across many tasks. However, if these syst…