works on

From the 1 of 12 linked papers with an AI index.

activity
20242026
collaborators

12 papers

cs.LG2026

HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models

Aznaur Aliev, Carlos Hinojosa, Abdelrahman Eldesokey +3

HyperSafe introduces a post‑hoc, model‑specific safe side network generated by a hypernetwork that classifies prompts using activation fingerprints, allowing fine‑tuned language mo…

cs.CR2026

Defending Against Harmful Supervision Hidden in Benign Samples

Bang An, Yibo Yang, Dandan Guo +3

Existing defenses are effective when harmful content is explicitly mixed into downstream fine-tuning data, but crafted samples can instead hide harmful supervision inside benign ta…

cs.CV2026

Mamba-FSCIL: Dynamic Adaptation with Selective State Space Model for Few-Shot Class-Incremental Learning

Xiaojie Li, Yibo Yang, Jianlong Wu +4

Few-shot class-incremental learning (FSCIL) aims to incrementally learn novel classes from limited examples while preserving knowledge of previously learned classes. Existing metho…

cs.AI2025

A Guardrail for Safety Preservation: When Safety-Sensitive Subspace Meets Harmful-Resistant Null-Space

Bingjie Zhang, Yibo Yang, Zhe Ren +4

Large language models (LLMs) have achieved remarkable success in diverse tasks, yet their safety alignment remains fragile during adaptation. Even when fine-tuning on benign data o…

cs.AI2025

Purifying Task Vectors in Knowledge-Aware Subspace for Model Merging

Bang An, Yibo Yang, Philip Torr +1

Model merging aims to integrate task-specific abilities from individually fine-tuned models into a single model without extra training. In recent model merging methods, task vector…

cs.LG2025

Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient

Zhongzhu Zhou, Yibo Yang, Ziyan Chen +7

Policy gradient (PG) methods in reinforcement learning frequently utilize deep neural networks (DNNs) to learn a shared backbone of feature representations used to compute likeliho…