Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Scalable Token-Level Hallucination Detection in Large Language Models
Rui Min, Tianyu Pang, Chao Du +2
Large language models (LLMs) have demonstrated remarkable capabilities, but they still frequently produce hallucinations. These hallucinations are difficult to detect in reasoning-…
cs.CL2025
Imperceptible Jailbreaking against Large Language Models
Kuofeng Gao, Yiming Li, Chao Du +4
Jailbreaking attacks on the vision modality typically rely on imperceptible adversarial perturbations, whereas attacks on the textual modality are generally assumed to require visi…