activity
20242026
collaborators

7 papers

cs.CL2026

Gavel: Agent Meets Checklist for Evaluating LLMs on Long-Context Legal Summarization

Yao Dou, Benjamin Mamut, Wei Xu

Large language models (LLMs) now support contexts of up to 1M tokens, but their strengths and weaknesses on complex long-context tasks remain unclear. To study this, we focus on mu…

cs.CL2026

Localizing Prompt Ambiguity in Large Language Models with Probe-Targeted Attribution

Govind Ramesh, Yao Dou, Wei Xu

Prompt ambiguity is a common source of failure in large language models, but is difficult to localize because it is a latent property of the prompt, while existing attribution meth…

cs.CL2025

SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants?

Yao Dou, Michel Galley, Baolin Peng +6

Large language models (LLMs) are increasingly used in interactive applications, and human evaluation remains the gold standard for assessing their performance in multi-turn convers…

cs.HC2024

Measuring, Modeling, and Helping People Account for Privacy Risks in Online Self-Disclosures with AI

Isadora Krsek, Anubha Kabra, Yao Dou +5

In pseudonymous online fora like Reddit, the benefits of self-disclosure are often apparent to users (e.g., I can vent about my in-laws to understanding strangers), but the privacy…

cs.CL2024

Towards Reliable Detection of LLM-Generated Texts: A Comprehensive Evaluation Framework with CUDRT

Zhen Tao, Yanfang Chen, Dinghao Xi +2

The increasing prevalence of large language models (LLMs) has significantly advanced text generation, but the human-like quality of LLM outputs presents major challenges in reliabl…

cs.CR2024

GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation

Govind Ramesh, Yao Dou, Wei Xu

Research on jailbreaking has been valuable for testing and understanding the safety and security issues of large language models (LLMs). In this paper, we introduce Iterative Refin…