5 papers
Learning Can Converge Stably to the Wrong Belief under Latent Reliability
Zhipeng Zhang, Zhenjie Yao, Kai Li +1
Learning systems are typically optimized by minimizing loss or maximizing reward, assuming that improvements in these signals reflect progress toward the true objective. However, w…
Meta-Cognitive Reinforcement Learning with Self-Doubt and Recovery
Zhipeng Zhang, Xiongfei Su, Kai Li
Robust reinforcement learning methods typically focus on suppressing unreliable experiences or corrupted rewards, but they lack the ability to reason about the reliability of their…
Stable but Wrong: When More Data Degrades Scientific Conclusions
Zhipeng Zhang, Kai Li
Modern science increasingly relies on ever-growing observational datasets and automated inference pipelines, under the implicit belief that accumulating more data makes scientific…
Learning to Trust Experience: A Monitor-Trust-Regulator Framework for Learning under Unobservable Feedback Reliability
Zhipeng Zhang, Zhenjie Yao, Kai Li +1
Learning under unobservable feedback reliability poses a distinct challenge beyond optimization robustness: a system must decide whether to learn from an experience, not only how t…
Training instability in deep learning follows low-dimensional dynamical principles
Zhipeng Zhang, Zhenjie Yao, Kai Li +1
Deep learning systems achieve remarkable empirical performance, yet the stability of the training process itself remains poorly understood. Training unfolds as a high-dimensional d…