2 papers
cs.AI2026
Dynamic Important Example Mining for Reinforcement Finetuning
Haoru Tan, Sitong Wu, Yanfeng Chen +9
Reinforcement fine-tuning (RFT) is increasingly used to strengthen the reasoning abilities of large models, yet its effectiveness is bound by how training data are selected and use…
cs.CV2026
Dataset Distillation by Influence Matching
Haoru Tan, Wang Wang, Sitong Wu +5
We revisit dataset distillation from an outcome-centric perspective. Rather than aligning process surrogates (per-step gradients or training trajectories), Influence Matching (Inf-…