collaborators

5 papers

cs.LG2026

Towards On-Policy SFT: Distribution Discriminant Theory and its Applications in LLM Training

Miaosen Zhang, Yishan Liu, Shuxia Lin +8

Supervised fine-tuning (SFT) is computationally efficient but often yields inferior generalization compared to reinforcement learning (RL). This gap is primarily driven by RL's use…

cs.CV2025

CONCORD: Concept-Informed Diffusion for Dataset Distillation

Jianyang Gu, Haonan Wang, Ruoxi Jia +4

Dataset distillation (DD) has witnessed significant progress in creating small datasets that encapsulate rich information from large original ones. Particularly, methods based on g…

cs.CL2025

Improved Methods for Model Pruning and Knowledge Distillation

Wei Jiang, Anying Fu, Youling Zhang

Model pruning is a performance optimization technique for large language models like R1 or o3-mini. However, existing pruning methods often lead to significant performance degradat…

cs.IR2025

ChorusCVR: Chorus Supervision for Entire Space Post-Click Conversion Rate Modeling

Wei Cheng, Yucheng Lu, Boyang Xia +9

Post-click conversion rate (CVR) estimation is a vital task in many recommender systems of revenue businesses, e.g., e-commerce and advertising. In a perspective of sample, a typic…

cs.LG2025

Group Distributionally Robust Dataset Distillation with Risk Minimization

Saeed Vahidian, Mingyu Wang, Jianyang Gu +3

Dataset distillation (DD) has emerged as a widely adopted technique for crafting a synthetic dataset that captures the essential information of a training dataset, facilitating the…