works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.CL2026

Can a Language Model Learn Facts Continually in Its Weights?

Charles O'Neill

The paper investigates whether language models can continuously acquire and retain factual knowledge directly in their weights, examining how different training styles affect reten…

cs.LG2026

DiscoGen: Procedural Generation of Algorithm Discovery Tasks in Machine Learning

Alexander D. Goldie, Zilin Wang, Adrian Hayler +17

Automating the development of machine learning algorithms has the potential to unlock new breakthroughs. However, our ability to improve and evaluate algorithm discovery systems ha…

cs.LG2026

Still: Amortized KV Cache Compaction in a Single Forward Pass

Charles O'Neill, Alex Sandomirsky, Harry Partridge +2

The KV cache is the memory bottleneck of long-horizon language model deployment. Practically, a deployable compactor must be lightweight enough to call during inference, expressive…

cs.LG2025

Resurrecting the Salmon: Rethinking Mechanistic Interpretability with Domain-Specific Sparse Autoencoders

Charles O'Neill, Mudith Jayasekara, Max Kirkby

Sparse autoencoders (SAEs) decompose large language model (LLM) activations into latent features that reveal mechanistic structure. Conventional SAEs train on broad data distributi…

cs.LG2025

A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations

Charles O'Neill, Slava Chalnev, Chi Chi Zhao +2

Contextual hallucinations -- statements unsupported by given context -- remain a significant challenge in AI. We demonstrate a practical interpretability insight: a generator-agnos…