3 papers
cs.CY2026
Understanding Content Moderation in Large Language Models through Restricted Books: From Refusal to Warning
Xucheng Yu, Emily Knox, Haohan Wang
As large language models enter everyday information pipelines, understanding how they handle sensitive topics matters as much as understanding whether they handle them at all. We s…
cs.LG2026
SCI-Defense: Defending Manipulation Attacks from Generative Engine Optimization
Xucheng Yu, Haibo Jin, Huimin Zeng +1
LLM-based ranking systems are vulnerable to Generative Engine Optimization (GEO) attacks, where adversaries inject semantic signals into product descriptions to artificially boost…
cs.AI2026
MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models
Han Wang, Yifan Sun, Brian Ko +8
Large language models (LLMs) can generate chains of thought (CoTs) that are not always causally responsible for their final outputs. When such a mismatch occurs, the CoT no longer…