works on

From the 2 of 5 linked papers with an AI index.

activity
20242026
collaborators

5 papers

cs.CL2026

Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response Rewriting

Mengya Hu, Susie Park, Suzana Ilic +5

The paper studies how to best place and combine content‑moderation filters and response rewriting in conversational systems, measuring overall usefulness and harmful exposure rathe…

cs.CL2026

From Prompt Risk to Response Risk: Paired Analysis of Safety Behavior of Large Language Models

Mengya Hu, Qiong Wei, Sandeep Atluri

The paper proposes a paired analysis framework that compares the risk level of prompts and their LLM-generated responses across multiple harm categories and severity levels, reveal…

cs.CL2025

FactCG: Enhancing Fact Checkers with Graph-Based Multi-Hop Data

Deren Lei, Yaxi Li, Siyao Li +6

Prior research on training grounded factuality classification models to detect hallucinations in large language models (LLMs) has relied on public natural language inference (NLI)…

cs.CR2024

InvisMark: Invisible and Robust Watermarking for AI-generated Image Provenance

Rui Xu, Mengya Hu, Deren Lei +6

The proliferation of AI-generated images has intensified the need for robust content authentication methods. We present InvisMark, a novel watermarking technique designed for high-…

cs.CL2024

SLM Meets LLM: Balancing Latency, Interpretability and Consistency in Hallucination Detection

Mengya Hu, Rui Xu, Deren Lei +5

Large language models (LLMs) are highly capable but face latency challenges in real-time applications, such as conducting online hallucination detection. To overcome this issue, we…