43 citations · 61 across the 39 of their papers we have counts for
59 papers
Learning to Follow In-Context Watermark Instructions via Self-Distillation
Yepeng Liu, Tianyi Chen, Xuandong Zhao +2
In-context watermarking (ICW) prepends an instruction to a query asking the model to embed a statistically detectable signal in its response. It thus equips LLMs with a watermarkin…
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
Xiaomin Li, Yuexing Hao, Jianheng Hou +90
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and inter…
When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study
Xiaolong Jin, Xuandong Zhao, Wenbo Guo +2
Reinforcement learning with verifiable rewards (RLVR) has driven substantial progress in large language model reasoning, but relies on ground-truth supervision that is costly or in…
VIMPO: Value-Implicit Policy Optimization for LLMs
Zhewei Kang, Aosong Feng, Sergey Levine +2
Reinforcement learning with verifiable rewards has become a central tool for improving the reasoning ability of large language models, but current methods face a trade-off between…
Reward Hacking in Language Model Agents: Revisiting AI Safety Gridworlds
Ömer Veysel Çağatan, Xuandong Zhao
Reward hacking, where AI systems exploit misspecified objectives to achieve high reward without satisfying intended goals, remains a central challenge in AI safety. Yet most known…
Audio Pirates: Black-box Audio Watermark Removal via Diffusion Priors
Lingfeng Yao, Xincong Zhong, Chenpei Huang +6
With the rise of AI-generated audio, watermarking has become widely used for detecting misuse and protecting intellectual property. However, adversaries may try to remove these wat…