4 papers
HarmProfile: Characterizing Harmful Distributions in Frontier LLMs
Zhouyuan Ma, Yutao Wu, Hanxun Huang +6
Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome rather than as an object of analysis. Consequently, little is kn…
Internal Safety Collapse in Frontier Large Language Models
Yutao Wu, Xiao Liu, Yifeng Gao +7
This work identifies a critical failure mode in frontier large language models (LLMs), which we term Internal Safety Collapse (ISC): under certain task conditions, models enter a s…
PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
Yutao Wu, Xiao Liu, Yunhao Feng +2
Large Language Models (LLMs) increasingly serve as research assistants, yet their reliability in scholarly tasks remains under-evaluated. In this work, we introduce PaperAsk, a ben…
ADMIT: Few-shot Knowledge Poisoning Attacks on RAG-based Fact Checking
Yutao Wu, Xiao Liu, Yinghui Li +5
Knowledge poisoning poses a critical threat to Retrieval-Augmented Generation (RAG) systems by injecting adversarial content into knowledge bases, tricking Large Language Models (L…