Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Calling the Bluff: Detecting Ever-Shifting Harmful Chat Dialogue via Ordered Reasoning Chain Regularization
Haojie Yu, Ziyou Jiang, Junjie Wang +4
Harmful chat dialogues are ever-shifting through type-shifting and lexical evasion, yet we find they share invariant principles, i.e., an Ordered Reasoning Chain (ORC) of recurring…
cs.CL2026
MAIGO: Mitigating Lost-in-Conversation with History-Cleaned On-Policy Self-Distillation
Haoyu Zheng, Yun Zhu, Shu Yuan +5
Large language models often solve tasks from a fully specified prompt but degrade when the same requirements unfold over multiple turns, known as the lost-in-conversation (LiC) gap…