Showing cs.CRShow all
3 papers · 1 filter
cs.CR2026
Layerwise Convergence Fingerprints for Runtime Misbehavior Detection in Large Language Models
Nay Myat Min, Long H. Pham, Jun Sun
Large language models deployed at runtime can misbehave in ways that clean-data validation cannot anticipate: training-time backdoors lie dormant until triggered, jailbreaks subver…
cs.CR2026
CORVUS: Red-Teaming Hallucination Detectors via Internal Signal Camouflage in Large Language Models
Nay Myat Min, Long H. Pham, Hongyu Zhang +1
Single-pass hallucination detectors rely on internal telemetry (e.g., uncertainty, hidden-state geometry, and attention) of large language models, implicitly assuming hallucination…
cs.CR2025
Unified Neural Backdoor Removal with Only Few Clean Samples through Unlearning and Relearning
Nay Myat Min, Long H. Pham, Jun Sun
Deep neural networks have achieved remarkable success across various applications; however, their vulnerability to backdoor attacks poses severe security risks -- especially in sit…