activity
20242026
collaborators

5 papers

cs.CL2026

Faster but Different: Diagnosing and Controlling Content Drift in Accelerated Multimodal Diffusion Language Models

Yaoxuan Dou, Yang Shu

Training-free acceleration makes diffusion-based multimodal large language models (dMLLMs) more deployable, but it may silently change generated content. We study this serving-time…

cs.CL2026

Gavel: Agent Meets Checklist for Evaluating LLMs on Long-Context Legal Summarization

Yao Dou, Benjamin Mamut, Wei Xu

Large language models (LLMs) now support contexts of up to 1M tokens, but their strengths and weaknesses on complex long-context tasks remain unclear. To study this, we focus on mu…

cs.CL2026

Localizing Prompt Ambiguity in Large Language Models with Probe-Targeted Attribution

Govind Ramesh, Yao Dou, Wei Xu

Prompt ambiguity is a common source of failure in large language models, but is difficult to localize because it is a latent property of the prompt, while existing attribution meth…

cs.CL2025

Evaluating LLMs on Chinese Idiom Translation

Cai Yang, Yao Dou, David Heineman +2

Idioms, whose figurative meanings usually differ from their literal interpretations, are common in everyday language, especially in Chinese, where they often contain historical ref…

cs.HC2024

Measuring, Modeling, and Helping People Account for Privacy Risks in Online Self-Disclosures with AI

Isadora Krsek, Anubha Kabra, Yao Dou +5

In pseudonymous online fora like Reddit, the benefits of self-disclosure are often apparent to users (e.g., I can vent about my in-laws to understanding strangers), but the privacy…