Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning
Yongcan Yu, Lingxiao He, Jian Liang +5
Test-time reinforcement learning (TTRL) always adapts models at inference time via pseudo-labeling, leaving it vulnerable to spurious optimization signals from label noise. Through…
cs.LG2024
Path-Specific Causal Reasoning for Fairness-aware Cognitive Diagnosis
Dacao Zhang, Kun Zhang, Le Wu +3
Cognitive Diagnosis~(CD), which leverages students and exercise data to predict students' proficiency levels on different knowledge concepts, is one of fundamental components in In…