From the 2 of 5 linked papers with an AI index.
5 papers
Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response Rewriting
Mengya Hu, Susie Park, Suzana Ilic +5
The paper studies how to best place and combine content‑moderation filters and response rewriting in conversational systems, measuring overall usefulness and harmful exposure rathe…
From Prompt Risk to Response Risk: Paired Analysis of Safety Behavior of Large Language Models
Mengya Hu, Qiong Wei, Sandeep Atluri
The paper proposes a paired analysis framework that compares the risk level of prompts and their LLM-generated responses across multiple harm categories and severity levels, reveal…
FactCG: Enhancing Fact Checkers with Graph-Based Multi-Hop Data
Deren Lei, Yaxi Li, Siyao Li +6
Prior research on training grounded factuality classification models to detect hallucinations in large language models (LLMs) has relied on public natural language inference (NLI)…
InvisMark: Invisible and Robust Watermarking for AI-generated Image Provenance
Rui Xu, Mengya Hu, Deren Lei +6
The proliferation of AI-generated images has intensified the need for robust content authentication methods. We present InvisMark, a novel watermarking technique designed for high-…
SLM Meets LLM: Balancing Latency, Interpretability and Consistency in Hallucination Detection
Mengya Hu, Rui Xu, Deren Lei +5
Large language models (LLMs) are highly capable but face latency challenges in real-time applications, such as conducting online hallucination detection. To overcome this issue, we…