activity
20242026
collaborators

5 papers

cs.LG2026

Position: Capability Control Should be a Separate Goal From Alignment

Shoaib Ahmed Siddiqui, Eleni Triantafillou, David Krueger +1

Foundation models are trained on broad data distributions, yielding generalist capabilities that enable many downstream applications but also expand the space of potential misuse a…

cs.LG2026

From Dormant to Deleted: Tamper-Resistant Unlearning Through Weight-Space Regularization

Shoaib Ahmed Siddiqui, Adrian Weller, David Krueger +3

Recent unlearning methods for LLMs are vulnerable to relearning attacks: knowledge believed-to-be-unlearned re-emerges by fine-tuning on a small set of (even seemingly-unrelated) e…

cs.LG2025

Neural Mutual Information Estimation with Vector Copulas

Yanzhi Chen, Zijing Ou, Adrian Weller +1

Estimating mutual information (MI) is a fundamental task in data science and machine learning. Existing estimators mainly rely on either highly flexible models (e.g., neural networ…

cs.CR2024

Countering Autonomous Cyber Threats

Kade M. Heckel, Adrian Weller

With the capability to write convincing and fluent natural language and generate code, Foundation Models present dual-use concerns broadly and within the cyber domain specifically.…

cs.LG2024

On Evaluating LLMs' Capabilities as Functional Approximators: A Bayesian Perspective

Shoaib Ahmed Siddiqui, Yanzhi Chen, Juyeon Heo +2

Recent works have successfully applied Large Language Models (LLMs) to function modeling tasks. However, the reasons behind this success remain unclear. In this work, we propose a…