2 papers
cs.LG2026
Adaptive Problem Generation via Symbolic Representations
Teresa Yeo, Myeongho Jeon, Dulaj Weerakoon +4
We present a method for generating training data for reinforcement learning with verifiable rewards to improve small open-weights language models on mathematical tasks. Existing da…
cs.LG2025
Test-Time Adaptation by Causal Trimming
Yingnan Liu, Rui Qiao, Mong Li Lee +1
Test-time adaptation aims to improve model robustness under distribution shifts by adapting models with access to unlabeled target samples. A primary cause of performance degradati…