13 papers
Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic Audio
Georgios Milis, Yubin Qin, Yihan Wu +1
As policy catches up with the capabilities of generative AI, watermarking is central to content provenance efforts. Inference-time watermarks for autoregressive models are unfit fo…
MCMark: Distortion-Free Multi-Bit Watermarking for Long Messages
Xuehao Cui, Ruibo Chen, Yihan Wu +1
Large language models now produce text indistinguishable from human writing, which increases the need for reliable provenance tracing. Multi-bit watermarking can embed identifiers…
More Haste, Less Speed: Weaker Single-Layer Watermark Improves Distortion-Free Watermark Ensembles
Ruibo Chen, Yihan Wu, Xuehao Cui +2
Watermarking has emerged as a crucial technique for detecting and attributing content generated by large language models. While recent advancements have utilized watermark ensemble…
Robust Distortion-Free Watermark for Autoregressive Audio Generation Models
Yihan Wu, Georgios Milis, Ruibo Chen +1
The rapid advancement of next-token-prediction models has led to widespread adoption across modalities, enabling the creation of realistic synthetic media. In the audio domain, whi…
Modality-Balancing Preference Optimization of Large Multimodal Models by Adversarial Negative Mining
Chenxi Liu, Tianyi Xiong, Yanshuo Chen +5
The task adaptation and alignment of Large Multimodal Models (LMMs) have been significantly advanced by instruction tuning and further strengthened by recent preference optimizatio…
Model Correlation Detection via Random Selection Probing
Ruibo Chen, Sheng Zhang, Yihan Wu +3
The growing prevalence of large language models (LLMs) and vision-language models (VLMs) has heightened the need for reliable techniques to determine whether a model has been fine-…