3 papers
cs.LG2026
Model Merging by Output-Space Projection
Bethan Evans, Benjamin Etheridge, Stephen Roberts +1
Model merging combines fine-tuned checkpoints into a single multi-task model without retraining. Existing methods - such as task arithmetic, model soups, TIES, and DARE - are compu…
cs.LG2025
Pruning Cannot Hurt Robustness: Certified Trade-offs in Reinforcement Learning
James Pedley, Benjamin Etheridge, Stephen J. Roberts +1
Reinforcement learning (RL) policies deployed in real-world environments must remain reliable under adversarial perturbations. At the same time, modern deep RL agents are heavily o…
cs.CL2025
Where Do Reasoning Models Refuse?
Kureha Yamaguchi, Benjamin Etheridge, Andy Arditi
Chat models without chain-of-thought (CoT) reasoning must decide whether to refuse a harmful request before generating their first response token. Reasoning models, by contrast, pr…