collaborators

5 papers

cs.LG2026

Position: Capability Control Should be a Separate Goal From Alignment

Shoaib Ahmed Siddiqui, Eleni Triantafillou, David Krueger +1

Foundation models are trained on broad data distributions, yielding generalist capabilities that enable many downstream applications but also expand the space of potential misuse a…

cs.LG2026

From Dormant to Deleted: Tamper-Resistant Unlearning Through Weight-Space Regularization

Shoaib Ahmed Siddiqui, Adrian Weller, David Krueger +3

Recent unlearning methods for LLMs are vulnerable to relearning attacks: knowledge believed-to-be-unlearned re-emerges by fine-tuning on a small set of (even seemingly-unrelated) e…

cs.LG2026

Permissive Information-Flow Analysis for Large Language Models

Shoaib Ahmed Siddiqui, Radhika Gaonkar, Boris Köpf +7

Large Language Models (LLMs) are rapidly becoming commodity components of larger software systems. This poses natural security and privacy problems: poisoned data retrieved from on…

cs.LG2024

Integrating uncertainty quantification into randomized smoothing based robustness guarantees

Sina Däubener, Kira Maag, David Krueger +1

Deep neural networks have proven to be extremely powerful, however, they are also vulnerable to adversarial attacks which can cause hazardous incorrect predictions in safety-critic…

cs.LG2024

Exploring the design space of deep-learning-based weather forecasting systems

Shoaib Ahmed Siddiqui, Jean Kossaifi, Boris Bonev +4

Despite tremendous progress in developing deep-learning-based weather forecasting systems, their design space, including the impact of different design choices, is yet to be well u…