1 citations · 1 across the 12 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
DenMark: Robust Semantic Watermarking for Diffusion Language Models
Tianhao Ma, Weihao Xuan, Dong-Dong Wu +5
Semantic text watermarks encode signals in meaning rather than surface token choices, offering robustness to paraphrasing and other semantic-preserving edits. Existing semantic wat…
cs.CL2026
The Confidence Dichotomy: Analyzing and Mitigating Miscalibration in Tool-Use Agents
Weihao Xuan, Qingcheng Zeng, Heli Qi +3
Autonomous agents based on large language models (LLMs) are rapidly evolving to handle multi-turn tasks, but ensuring their trustworthiness remains a critical challenge. A fundamen…
cs.CL2025
MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation
Weihao Xuan, Rui Yang, Heli Qi +29
Existing large language model (LLM) evaluation benchmarks primarily focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-lingui…