3 papers
cs.LG2026
PowerStep: Memory-Efficient Adaptive Optimization via -Norm Steepest Descent
Yao Lu, Dengdong Fan, Shixun Zhang +1
Adaptive optimizers, most notably Adam, have become the default standard for training large-scale neural networks such as Transformers. These methods maintain running estimates of…
physics.comp-ph2026
SMC-AI: Scaling Monte Carlo Simulation to Four Trillion Atoms with AI Accelerators
Xianglin Liu, Kai Yang, Fanli Zhou +9
The rapid advancement of deep learning is reshaping the hardware design landscape toward AI tasks, posing fundamental challenges for HPC workloads such as atomistic simulation. Her…
cs.LG2025
Calibrating and Rotating: A Unified Framework for Weight Conditioning in PEFT
Da Chang, Peng Xue, Yu Li +3
Parameter-Efficient Fine-Tuning (PEFT) methods are crucial for adapting large pre-trained models. Among these, LoRA is considered a foundational approach. Building on this, the inf…