1 citations · 1 across the 6 of their papers we have counts for
4 papers · 1 filter
RepSelect: Robust LLM Unlearning via Representation Selectivity
Filip Sondej, Yushi Yang, Adam Mahdi
When LLM weights are open or fine-tuning is available through an API, suppressing hazardous knowledge and tendencies is not enough: removal has to be deep enough that an adversary…
Agentic Reinforcement Learning for Search Misaligns Instruction-Tuning
Yushi Yang, Shreyansh Padarha, Sarah Ball +2
Agentic reinforcement learning (RL) trains large language models to use tools, but its impact on alignment is poorly understood. We study how agentic RL for search affects the alig…
Measuring what Matters: Construct Validity in Large Language Model Benchmarks
Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou +39
Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstrac…
Evaluating Fine-Tuning Efficiency of Human-Inspired Learning Strategies in Medical Question Answering
Yushi Yang, Andrew M. Bean, Robert McCraith +1
Fine-tuning Large Language Models (LLMs) incurs considerable training costs, driving the need for data-efficient training with optimised data ordering. Human-inspired strategies of…