collaborators

6 papers

cs.LG2026

Decoupling the "What" and "Where" With Polar Coordinate Positional Embeddings

Anand Gopalakrishnan, Robert Csordás, Jürgen Schmidhuber +1

The attention mechanism in a Transformer architecture matches key to query based on both content -- the what -- and position in a sequence -- the where. We present an analysis indi…

cs.CL2026

Catch Your Breath: Adaptive Computation for Self-Paced Sequence Production

Alexandre Galashov, Matt Jones, Rosemary Ke +3

Within the landscape of inference-time scaling methods for foundation models, a width-based approach to scaling -- which involves the insertion of <pause> tokens in the input strea…

cs.LG2026

Is your algorithm unlearning or untraining?

Eleni Triantafillou, Ahmed Imtiaz Humayun, Monica Ribero +3

As models are getting larger and are trained on increasing amounts of data, there has been an explosion of interest into how we can ``delete'' specific data points or behaviours fr…

cs.LG2026

Improving Discrete Optimisation Via Decoupled Straight-Through Estimator

Rushi Shah, Mingyuan Yan, Michael Curtis Mozer +1

The Straight-Through Estimator (STE) is the dominant method for training neural networks with discrete variables, enabling gradient-based optimisation by routing gradients through…

cs.LG2026

Redirection for Erasing Memory (REM): Towards a universal unlearning method for corrupted data

Stefan Schoepf, Michael Curtis Mozer, Nicole Elyse Mitchell +4

Machine unlearning is studied for a multitude of tasks, but specialization of unlearning methods to particular tasks has made their systematic comparison challenging. To address th…

cs.LG2026

From Dormant to Deleted: Tamper-Resistant Unlearning Through Weight-Space Regularization

Shoaib Ahmed Siddiqui, Adrian Weller, David Krueger +3

Recent unlearning methods for LLMs are vulnerable to relearning attacks: knowledge believed-to-be-unlearned re-emerges by fine-tuning on a small set of (even seemingly-unrelated) e…