2 papers
cs.AI2026
Adaptive Teacher Exposure for Self-Distillation in LLM Reasoning
Zihao Han, Tiangang Zhang, Huaibin Wang +1
On-policy self-distillation has become a strong recipe for LLM reasoning, where a privileged teacher supervises the student's own rollouts while conditioning on the reference solut…
cs.LG2026
Reinforcement Learning with Markov Risk Measures and Multipattern Risk Approximation
Andrzej Ruszczynski, Tiangang Zhang
For a risk-averse finite-horizon Markov Decision Problem, we introduce a special class of Markov coherent risk measures, called mini-batch measures. We also define the class of mul…