Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
GATES: Self-Distillation under Privileged Context with Consensus Gating
Alex Stein, Furong Huang, Tom Goldstein
We study self-distillation in settings where supervision is unreliable: there are no ground truth labels, verifiable rewards, or external graders to evaluate answers. We focus on d…
cs.LG2026
Provably Efficient Algorithms for S- and Non-Rectangular Robust MDPs with General Parameterization
Anirudh Satheesh, Ziyi Chen, Furong Huang +1
We study robust Markov decision processes (RMDPs) with general policy parameterization under s-rectangular and non-rectangular uncertainty sets. Prior work is largely limited to ta…