5 papers
From Curiosity to Caution: Mitigating Reward Hacking for Best-of-N with Pessimism
Zhuohao Yu, Zhiwei Steven Wu, Adam Block
Inference-time compute scaling has emerged as a powerful paradigm for improving language model performance on a wide range of tasks, but the question of how best to use the additio…
Partition Function Estimation under Bounded f-Divergence
Adam Block, Abhishek Shetty
We study the statistical complexity of estimating partition functions given sample access to a proposal distribution and an unnormalized density ratio for a target distribution. Wh…
MarkTune: Improving the Quality-Detectability Trade-off in Open-Weight LLM Watermarking
Yizhou Zhao, Zhiwei Steven Wu, Adam Block
Watermarking aims to embed hidden signals in generated text that can be reliably detected when given access to a secret key. Open-weight language models pose acute challenges for s…
Small Loss Bounds for Online Learning Separated Function Classes: A Gaussian Process Perspective
Adam Block, Abhishek Shetty
In order to develop practical and efficient algorithms while circumventing overly pessimistic computational lower bounds, recent work has been interested in developing oracle-effic…
GaussMark: A Practical Approach for Structural Watermarking of Language Models
Adam Block, Ayush Sekhari, Alexander Rakhlin
Recent advances in Large Language Models (LLMs) have led to significant improvements in natural language processing tasks, but their ability to generate human-quality text raises s…