3 papers
cs.LG2025
Provable Benefit of Sign Descent: A Minimal Model Under Heavy-Tailed Class Imbalance
Robin Yadav, Shuo Xie, Tianhao Wang +1
Adaptive optimization methods (such as Adam) play a major role in LLM pretraining, significantly outperforming Gradient Descent (GD). Recent studies have proposed new smoothness as…
cs.LG2025
A Tale of Two Geometries: Adaptive Optimizers and Non-Euclidean Descent
Shuo Xie, Tianhao Wang, Beining Wu +1
Adaptive optimizers can reduce to normalized steepest descent (NSD) when only adapting to the current gradient, suggesting a close connection between the two algorithmic families.…
cs.LG2025
Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation
Mohamad Amin Mohamadi, Tianhao Wang, Zhiyuan Li
Modern language models fail a fundamental requirement of trustworthy intelligence: knowing when not to answer. Despite achieving impressive accuracy on benchmarks, these models pro…