8 papers
Randomized Midpoint Method for Log-Concave Sampling under Constraints
Yifeng Yu, Shijie Zhang, Lu Yu
In this paper, we study the problem of sampling from log-concave distributions supported on convex and compact sets, with a particular focus on the randomized midpoint discretizati…
On the Limits of Latent Reuse in Diffusion Models
Yifeng Yu, Lu Yu
Diffusion models are often trained in low-dimensional latent spaces, which are then reused for related but shifted datasets. In this work, we study when such latent reuse remains r…
Nonexistence of vanishing-viscosity limits for mechanical Hamiltonian ergodic problems
Ziran Liu, Hung V. Tran, Yifeng Yu
For , let be the solution of the ergodic problem \[ \frac12 |DÏ^\varepsilon|^2+F(x)-\varepsilonÎÏ^\varepsilon=c(\varepsilon) \qquad \text{on } \m…
BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning
Yuhang Xu, Kaibin Tian, Yang Tian +6
Reinforcement Learning (RL) has become a cornerstone for improving the performance of Large Language Models (LLMs). However, its rollout phase constitutes a significant efficiency…
InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and Repetition
Fengze Liu, Weidong Zhou, Binbin Liu +7
Upweighting high-quality data in LLM pretraining often improves performance, but in datalimited regimes, especially under overtraining, stronger upweighting increases repetition an…
Diffusion Models with Heavy-Tailed Targets: Score Estimation and Sampling Guarantees
Yifeng Yu, Lu Yu
Score-based diffusion models have become a powerful framework for generative modeling, with score estimation as a central statistical bottleneck. Existing guarantees for score esti…