4 papers
LLM Watermarking Using Mixtures and Statistical-to-Computational Gaps
Pedro Abdalla, Roman Vershynin
Given a text, can we determine whether it was generated by a large language model (LLM) or by a human? A widely studied approach to this problem is watermarking. We propose an unde…
Thinning to improve two-sample discrepancy
Gleb Smirnov, Roman Vershynin
The discrepancy between two independent samples \(X_1,\dots,X_n\) and \(Y_1,\dots,Y_n\) drawn from the same distribution on typically has order \(O(\sqrt{n})\) even…
On the Dimension-Free Concentration of Simple Tensors via Matrix Deviation
Pedro Abdalla, Roman Vershynin
We provide a simpler proof of a sharp concentration inequality for subgaussian simple tensors obtained recently by Al-Ghattas, Chen and Sanz-Alonso. Our approach uses a matrix devi…
Improving discrepancy by moving a few points
Gleb Smirnov, Roman Vershynin
We show how to improve the discrepancy of an iid sample by moving only a few points. Specifically, modifying \( O(m) \) sample points on average reduces the Kolmogorov-Smirnov dist…