7 papers · 1 filter
Epiphany-Aware KV Cache Eviction Without the Attention Matrix
Steven Kolawole, Virginia Smith
As reasoning models emit chains of thought tens of thousands of tokens long, KV cache increasingly becomes a deployment bottleneck. Existing cache eviction methods rank tokens by a…
What Survives When You Compress a Recursive Reasoner for the Edge?
Pearse Jim, Steven Kolawole, Opegbemi Matthias Busoye +2
Recursive reasoning models can solve complex structured tasks with only a few million parameters by repeatedly updating a latent state. Deploying these models on edge hardware requ…
Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks
Kevin Kuo, Chhavi Yadav, Virginia Smith
Recent defenses for safeguarding open-weight large language models (LLMs) are intended to prevent adversarial usage. Underlying these defenses is an assumption that new harmful beh…
Curriculum Learning for Safety Alignment
Sandeep Kumar, Virginia Smith, Chhavi Yadav
Direct Preference Optimisation (DPO) is widely used for safety alignment in large language models. However, prior work shows it is brittle and exhibits poor out-of-distribution (OO…
Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
Steven Kolawole, Lucio Dery, Jean-François Kagy +3
Structured pruning is a promising approach to create smaller, faster large language models. However, existing methods typically rely on computing the gradient via backward passes,…
PARALLELPROMPT: Extracting Parallelism from Large Language Model Queries
Steven Kolawole, Keshav Santhanam, Virginia Smith +1
LLM serving systems typically treat user prompts as monolithic inputs, optimizing inference through decoding tricks or inter-query batching. However, many real-world prompts contai…