content moderation 1filter placement 1harm severity assessment 1latency 1llm safety 1paired evaluation 1prompt-response analysis 1response rewriting 1risk escalation 1safety 1
From the 2 of 3 linked papers with an AI index.
3 papers
cs.CL2026
Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response Rewriting
Mengya Hu, Susie Park, Suzana Ilic +5
The paper studies how to best place and combine content‑moderation filters and response rewriting in conversational systems, measuring overall usefulness and harmful exposure rathe…
cs.CL2026
From Prompt Risk to Response Risk: Paired Analysis of Safety Behavior of Large Language Models
Mengya Hu, Qiong Wei, Sandeep Atluri
The paper proposes a paired analysis framework that compares the risk level of prompts and their LLM-generated responses across multiple harm categories and severity levels, reveal…
cs.CL2025
FactCG: Enhancing Fact Checkers with Graph-Based Multi-Hop Data
Deren Lei, Yaxi Li, Siyao Li +6
Prior research on training grounded factuality classification models to detect hallucinations in large language models (LLMs) has relied on public natural language inference (NLI)…