5 citations · 5 across the 7 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
A Tale of Two Geometries: Adaptive Optimizers and Non-Euclidean Descent
Shuo Xie, Tianhao Wang, Beining Wu +1
Adaptive optimizers can reduce to normalized steepest descent (NSD) when only adapting to the current gradient, suggesting a close connection between the two algorithmic families.…
cs.LG2025
Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression
Xingwu Chen, Miao Lu, Beining Wu +1
Using more test-time computation during language model inference, such as generating more intermediate thoughts or sampling multiple candidate answers, has proven effective in sign…