2 papers
cs.CR2026
Activation-Guided Local Editing for Jailbreaking Attacks
Jiecong Wang, Haoran Li, Hao Peng +4
Jailbreaking is an essential adversarial technique for red-teaming these models to uncover and patch security flaws. However, existing jailbreak methods face significant drawbacks.…
cs.AI2026
Latent Chain-of-Thought as Planning: Decoupling Reasoning from Verbalization
Jiecong Wang, Hao Peng, Chunyang Liu
Chain-of-Thought (CoT) empowers Large Language Models (LLMs) to tackle complex problems, but remains constrained by the computational cost and reasoning path collapse when grounded…