activity
20202025
most citedTransferred Discrepancy: Quantifying the Difference Between Representations

6 citations · 6 across the 4 of their papers we have counts for

collaborators

7 papers

cs.LG2025

Don't Waste Mistakes: Leveraging Negative RL-Groups via Confidence Reweighting

Yunzhen Feng, Parag Jain, Anthony Hartshorn +2

Reinforcement learning with verifiable rewards (RLVR) has become a standard recipe for improving large language models (LLMs) on reasoning tasks, with Group Relative Policy Optimiz…

cs.LG2025

What Characterizes Effective Reasoning? Revisiting Length, Review, and Structure of CoT

Yunzhen Feng, Julia Kempe, Cheng Zhang +2

Large reasoning models (LRMs) spend substantial test-time compute on long chain-of-thought (CoT) traces, but what *characterizes* an effective CoT remains unclear. While prior work…

cs.LG2025

PILAF: Optimal Human Preference Sampling for Reward Modeling

Yunzhen Feng, Ariel Kwiatkowski, Kunhao Zheng +2

As large language models increasingly drive real-world applications, aligning them with human values becomes paramount. Reinforcement Learning from Human Feedback (RLHF) has emerge…

cs.LG2025

Spend Wisely: Maximizing Post-Training Gains in Iterative Synthetic Data Bootstrapping

Pu Yang, Yunzhen Feng, Ziyuan Chen +2

Modern foundation models often undergo iterative ``bootstrapping'' in their post-training phase: a model generates synthetic data, an external verifier filters out low-quality samp…

cs.LG2024

Strong Model Collapse

Elvis Dohmatob, Yunzhen Feng, Arjun Subramonian +1

Within the scaling laws paradigm, which underpins the training of large neural networks like ChatGPT and Llama, we consider a supervised regression setting and establish the exista…

cs.LG20206 cited

Transferred Discrepancy: Quantifying the Difference Between Representations

Yunzhen Feng, Runtian Zhai, Di He +2

Understanding what information neural networks capture is an essential problem in deep learning, and studying whether different models capture similar features is an initial step t…