2 citations · 5 across the 6 of their papers we have counts for
Showing 2026 · cs.CRShow all
2 papers · 2 filters
cs.CR2026
MechAudit-40: White-Box Auditing across 40 LLM Attack Mechanisms
Zhen Guo, Shanghao Shi, Shamim Yazdani +2
While LLM attacks span prompt optimization, multi-turn context manipulation, retrieval poisoning, and model backdoors, white-box defenses are typically evaluated on isolated attack…
cs.CR2026
TraceGuard: Process-Guided Firewall against Reasoning Backdoors in Large Language Models
Zhen Guo, Shanghao Shi, Hao Li +3
Large Reasoning Models (LRMs) introduce a reasoning-level attack surface: adversaries can corrupt intermediate inferences while preserving a plausible trace and an apparently benig…