4 papers
CORVUS: Red-Teaming Hallucination Detectors via Internal Signal Camouflage in Large Language Models
Nay Myat Min, Long H. Pham, Hongyu Zhang +1
Single-pass hallucination detectors rely on internal telemetry (e.g., uncertainty, hidden-state geometry, and attention) of large language models, implicitly assuming hallucination…
AUTOLAW: Enhancing Legal Compliance in Large Language Models via Case Law Generation and Jury-Inspired Deliberation
Tai D. Nguyen, Long H. Pham, Jun Sun
The rapid advancement of domain-specific large language models (LLMs) in fields like law necessitates frameworks that account for nuanced regional legal distinctions, which are cri…
Propaganda AI: An Analysis of Semantic Divergence in Large Language Models
Nay Myat Min, Long H. Pham, Yige Li +1
Large language models (LLMs) can exhibit concept-conditioned semantic divergence: common high-level cues (e.g., ideologies, public figures) elicit unusually uniform, stance-like re…
CROW: Eliminating Backdoors from Large Language Models via Internal Consistency Regularization
Nay Myat Min, Long H. Pham, Yige Li +1
Large Language Models (LLMs) are vulnerable to backdoor attacks that manipulate outputs via hidden triggers. Existing defense methods--designed for vision/text classification tasks…