3 papers
cs.LG2026
Model Merging by Output-Space Projection
Bethan Evans, Benjamin Etheridge, Stephen Roberts +1
Model merging combines fine-tuned checkpoints into a single multi-task model without retraining. Existing methods - such as task arithmetic, model soups, TIES, and DARE - are compu…
cs.CL2026
Where Do Reasoning Models Refuse?
Kureha Yamaguchi, Benjamin Etheridge, Andy Arditi
Chat models without chain-of-thought (CoT) reasoning must decide whether to refuse a harmful request before generating their first response token. Reasoning models, by contrast, pr…
cs.LG2025
Pruning Cannot Hurt Robustness: Certified Trade-offs in Reinforcement Learning
James Pedley, Benjamin Etheridge, Stephen J. Roberts +1
Reinforcement learning (RL) policies deployed in real-world environments must remain reliable under adversarial perturbations. At the same time, modern deep RL agents are heavily o…