3 papers
cs.LG2026
PowerStep: Memory-Efficient Adaptive Optimization via -Norm Steepest Descent
Yao Lu, Dengdong Fan, Shixun Zhang +1
Adaptive optimizers, most notably Adam, have become the default standard for training large-scale neural networks such as Transformers. These methods maintain running estimates of…
cs.LG2026
Sub-JEPA: Subspace Gaussian Regularization for Stable End-to-End World Models
Kai Zhao, Dongliang Nie, Yuchen Lin +4
Joint-Embedding Predictive Architectures (JEPAs) provide a simpleframework for learning world models by predicting future latent representations.However, JEPA training is subject t…
cs.CV2026
1%>100%: High-Efficiency Visual Adapter with Complex Linear Projection Optimization
Dongshuo Yin, Xue Yang, Deng-Ping Fan +1
Deploying vision foundation models typically relies on efficient adaptation strategies, whereas conventional full fine-tuning suffers from prohibitive costs and low efficiency. Whi…