From the 1 of 4 linked papers with an AI index.
4 papers
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models
Aznaur Aliev, Carlos Hinojosa, Abdelrahman Eldesokey +3
HyperSafe introduces a post‑hoc, model‑specific safe side network generated by a hypernetwork that classifies prompts using activation fingerprints, allowing fine‑tuned language mo…
Defending Against Harmful Supervision Hidden in Benign Samples
Bang An, Yibo Yang, Dandan Guo +3
Existing defenses are effective when harmful content is explicitly mixed into downstream fine-tuning data, but crafted samples can instead hide harmful supervision inside benign ta…
Purifying Task Vectors in Knowledge-Aware Subspace for Model Merging
Bang An, Yibo Yang, Philip Torr +1
Model merging aims to integrate task-specific abilities from individually fine-tuned models into a single model without extra training. In recent model merging methods, task vector…
Dynamic Context-oriented Decomposition for Task-aware Low-rank Adaptation with Less Forgetting and Faster Convergence
Yibo Yang, Sihao Liu, Chuan Rao +5
Conventional low-rank adaptation methods build adapters without considering data context, leading to sub-optimal fine-tuning performance and severe forgetting of inherent world kno…