3 papers
cs.CR2026
OTTER: A Red-Teaming System for Toxicity-Evading Jailbreak Prompt Optimization
Jerry Wang, Hsin-Ling Hsu, Yi-Cheng Lai +2
Production LLMs increasingly rely on toxicity-based moderation filters as a primary defense, assuming that harmful intent correlates with toxic surface wording. We show this assump…
cs.CL2026
MedAction: Towards Active Multi-turn Clinical Diagnostic LLMs
Hsin-Ling Hsu, Zizheng Wang, Donghua Zhang +9
Most existing LLM diagnoses are evaluated on static, single-turn settings where complete patient information is provided upfront, an oversimplification of real clinical practice. W…
cs.LG2026
WARP: Guaranteed Inner-Layer Repair of NLP Transformers
Hsin-Ling Hsu, Min-Yu Chen, Nai-Chia Chen +3
Transformer-based NLP models remain vulnerable to adversarial perturbations, yet existing repair methods face a fundamental trade-off: gradient-based approaches offer flexibility b…