5 papers
Multi-Mask Diffusion Language Models for Few-Step Generation
Sijin Chen, Yinuo Ren, Heyang Zhao +3
Masked diffusion models (MDMs) are a promising family of language generators, but achieving high-quality few-step generation remains challenging. In MDMs, all forward trajectories…
Relative Translation Invariant Wasserstein Distance
Binshuai Wang, Qiwei Di, Ming Yin +3
Motivated by the Bures distance, we introduce a new family of distances, \emph{relative translation invariant Wasserstein distances}, denoted by , as an extension of the clas…
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
Kaixuan Ji, Qiwei Di, Heyang Zhao +2
Kullback-Leibler (KL) regularization is widely used in offline decision-making and offers several benefits, motivating recent work on the sample complexity of offline learning with…
Near-Optimal Regret for KL-Regularized Multi-Armed Bandits
Kaixuan Ji, Qingyue Zhao, Heyang Zhao +2
Recent studies have shown that reinforcement learning with KL-regularized objectives can enjoy faster rates of convergence or logarithmic regret, in contrast to the classical $\sqr…
On the Limits of Test-Time Compute: Sequential Reward Filtering for Better Inference
Yue Yu, Qiwei Di, Quanquan Gu +1
Test-time compute (TTC) has become an increasingly prominent paradigm for enhancing large language models (LLMs). Despite the empirical success of methods such as best-of- (BoN)…