Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Beyond Hate: Differentiating Uncivil and Intolerant Speech in Multimodal Content Moderation
Nils A. Herrmann, Tobias Eder, Jingyi He +1
Current multimodal toxicity benchmarks typically use a single binary hatefulness label. This coarse approach conflates two fundamentally different characteristics of expression: to…
cs.CL2026
Safer Reasoning Traces: Measuring and Mitigating Chain-of-Thought Leakage in LLMs
Patrick Ahrend, Tobias Eder, Xiyang Yang +2
Chain-of-Thought (CoT) prompting improves LLM reasoning but can increase privacy risk by resurfacing personally identifiable information (PII) from the prompt into reasoning traces…