4 papers
PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks
Zhenxin Ai, Haiyun He
Watermarking for large language models (LLMs) is a promising approach for detecting LLM-generated text and enabling responsible deployment. However, existing watermarking methods a…
On the Generalization of Knowledge Distillation: An Information-Theoretic View
Bingying Li, Haiyun He
Knowledge distillation is widely used to improve generalization in practice, yet its theoretical understanding remains elusive. In the standard distillation setting, a teacher mode…
Fundamental Trade-Offs in Multi-Bit Watermarking of Stochastic Processes
Haiyun He, Yepeng Liu, Zhuoer Shen +3
We study multi-bit watermarking for data generated by stochastic processes, where a hidden message is embedded during sampling and must be decodable by an authorized detector that…
Information-Theoretic Generalization Bounds for Deep Neural Networks
Haiyun He, Ziv Goldfeld
Deep neural networks (DNNs) exhibit an exceptional capacity for generalization in practical applications. This work aims to capture the effect and benefits of depth for supervised…