2 papers
cs.CL2026
DECOR: Auditing LLM Deception via Information Manipulation Theory
Linyue Cai, Samuel Yeh, Jwala Dhamala +2
Large language models can deceive by subtly manipulating truthful information -- omitting key facts, shifting focus, or obscuring meaning -- making such behavior difficult to detec…
cs.AI2025
Certifying Counterfactual Bias in LLMs
Isha Chaudhary, Qian Hu, Manoj Kumar +3
Large Language Models (LLMs) can produce biased responses that can cause representational harms. However, conventional studies are insufficient to thoroughly evaluate biases across…