3 papers
cs.CR2025
Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race?
Yuan Xin, Dingfan Chen, Linyi Yang +2
As large language models (LLMs) are increasingly deployed, ensuring their safe use is paramount. Jailbreaking, adversarial prompts that bypass model alignment to trigger harmful ou…
cs.AI2025
ResearStudio: A Human-Intervenable Framework for Building Controllable Deep-Research Agents
Linyi Yang, Yixuan Weng
Current deep-research agents run in a ''fire-and-forget'' mode: once started, they give users no way to fix errors or add expert knowledge during execution. We present ResearStudio…
cs.LG2025
Anchored Supervised Fine-Tuning
He Zhu, Junyou Su, Peng Lai +4
Post-training of large language models involves a fundamental trade-off between supervised fine-tuning (SFT), which efficiently mimics demonstrations but tends to memorize, and rei…