9 citations · 28 across the 19 of their papers we have counts for
Showing 2026Show all
2 papers · 1 filter
cs.LG2026
TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment
Changyue Li, Jiaming He, Youliang Yuan +4
Fine-Tuning-as-a-Service (FTaaS) platforms let users train large language models (LLMs) on customized tasks, but this pipeline could erode models' safety alignment. In practice, se…
cs.CL2026
SHAPE: Unifying Safety, Helpfulness and Pedagogy for Educational LLMs
Sihang Zhao, Kangrui Yu, Youliang Yuan +2
Large Language Models (LLMs) have been widely explored in educational scenarios. We identify a critical vulnerability in current educational LLMs, pedagogical jailbreaks, where stu…