6 papers
Reinforcement Learning without Ground-Truth Solutions can Improve LLMs
Yingyu Lin, Qiyue Gao, Nikki Lijing Kuang +6
Reinforcement learning with verifiable rewards (RLVR) for training LLMs typically rely on ground-truth answers to assign rewards, limiting their applicability to tasks where the gr…
On the -Free Inference Complexity of Absorbing Discrete Diffusion
Xunpeng Huang, Yingyu Lin, Nishant Jain +4
Absorbing discrete diffusion has emerged as a dominant framework for discrete data generation. However, a significant disparity remains between its empirical success and theoretica…
Purifying Approximate Differential Privacy with Randomized Post-processing
Yingyu Lin, Erchi Wang, Yi-An Ma +1
We propose a framework to convert -approximate Differential Privacy (DP) mechanisms into -pure DP mechanisms under certain conditions, a proce…
Beyond Length: Quantifying Long-Range Information for Long-Context LLM Pretraining Data
Haoran Deng, Yingyu Lin, Zhenghao Lin +4
Long-context language models unlock advanced capabilities in reasoning, code generation, and document summarization by leveraging dependencies across extended spans of text. Howeve…
Almost Linear Convergence under Minimal Score Assumptions: Quantized Transition Diffusion
Xunpeng Huang, Yingyu Lin, Nikki Lijing Kuang +4
Continuous diffusion models have demonstrated remarkable performance in data generation across various domains, yet their efficiency remains constrained by two critical limitations…
A Skewness-Based Criterion for Addressing Heteroscedastic Noise in Causal Discovery
Yingyu Lin, Yuxing Huang, Wenqin Liu +6
Real-world data often violates the equal-variance assumption (homoscedasticity), making it essential to account for heteroscedastic noise in causal discovery. In this work, we expl…