1 citations · 2 across the 14 of their papers we have counts for
4 papers · 1 filter
Towards Context-Invariant Safety Alignment for Large Language Models
Yixu Wang, Yang Yao, Xin Wang +4
Preference-based post-training aligns LLMs with human intent, yet safety behavior often remains brittle. A model may refuse a harmful request in a standard prompt but comply when t…
Internal Safety Collapse in Frontier Large Language Models
Yutao Wu, Xiao Liu, Yifeng Gao +7
This work identifies a critical failure mode in frontier large language models (LLMs), which we term Internal Safety Collapse (ISC): under certain task conditions, models enter a s…
ADMIT: Few-shot Knowledge Poisoning Attacks on RAG-based Fact Checking
Yutao Wu, Xiao Liu, Yinghui Li +5
Knowledge poisoning poses a critical threat to Retrieval-Augmented Generation (RAG) systems by injecting adversarial content into knowledge bases, tricking Large Language Models (L…
Identity Lock: Locking API Fine-tuned LLMs With Identity-based Wake Words
Hongyu Su, Yifeng Gao, Yifan Ding +1
The rapid advancement of Large Language Models (LLMs) has increased the complexity and cost of fine-tuning, leading to the adoption of API-based fine-tuning as a simpler and more e…