Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Learning Human Health and Diseases from 24-hour Wrist Movement
Yong Wang, Dylan McGagh, Katya Broomberg +21
Much of human health and function unfolds beyond the clinic, through the movements of everyday life. Wrist-worn accelerometers capture these movements continuously, yet their rich…
cs.LG2026
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation
Yang Sun, Lichao Ma, Houyuan Qin +5
On-policy distillation (OPD) trains a student on its own trajectories under dense token-level supervision from a teacher. Reward-extrapolation methods such as ExOPD amplify the tea…