3 papers
cs.LG2026
Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization
Huilin Zhou, Jian Zhao, Yilu Zhong +7
Red teaming is critical for uncovering vulnerabilities in Large Language Models (LLMs). While automated methods have improved scalability, existing approaches often rely on static…
cs.AI2026
Nice Fold or Hero Call: Learning Budget-Efficient Thinking for Adaptive Reasoning
Zhaomeng Zhou, Lan Zhang, Junyang Wang +2
Large reasoning models (LRMs) improve problem solving through extended reasoning, but often misallocate test-time compute. Existing efficiency methods reduce cost by compressing re…
cs.CR2026
Permit: Permission-Aware Representation Intervention for Controlled Generation in Large Language Models
Pengcheng Sun, Lan Zhang, Zhaopeng Zhang +2
Large language models (LLMs) are increasingly deployed in enterprise settings where they handle sensitive documents and user context, raising acute concerns over security and contr…