17 papers
Optimal Watermark Localization in Mixed-Source Large Language Model Texts
Jose H. Blanchet, T. Tony Cai, Xiang Li +3
Watermarking provides a principled way to authenticate text generated by large language models (LLMs). In practice, however, the final text may be mixed-source, with watermark evid…
CurveRL: Principled Distribution-Aware Context Reweighting for LLM Reasoning
Ke Sun, Yizhou Zhao, Jiayi Xin +2
Context or prompt-level reweighting has emerged as a central algorithmic lever in Reinforcement Learning with Verified Rewards (RLVR) for improving the reasoning capability of larg…
Robust Spectral Watermark for Synthetic Tabular Data
Yizhou Zhao, Xiang Li, Peter Song +2
The rise of generative AI has enabled the production of high-fidelity synthetic tabular data across fields such as healthcare, finance, and public policy, raising growing concerns…
Statistical Impossibility and Possibility of Aligning LLMs with Human Preferences: From Condorcet Paradox to Nash Equilibrium
Kaizhao Liu, Qi Long, Zhekun Shi +2
Aligning large language models (LLMs) with diverse human preferences is critical for ensuring fairness and informed outcomes when deploying these models for decision-making. In thi…
Improving the Trade-off Between Watermark Strength and Speculative Sampling Efficiency for Language Models
Weiqing He, Xiang Li, Li Shen +2
Watermarking is a principled approach for tracing the provenance of large language model (LLM) outputs, but its deployment in practice is hindered by inference inefficiency. Specul…
On the Empirical Power of Goodness-of-Fit Tests in Watermark Detection
Weiqing He, Xiang Li, Tianqi Shang +3
Large language models (LLMs) raise concerns about content authenticity and integrity because they can generate human-like text at scale. Text watermarks, which embed detectable sta…