Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference
Shashank Kapadia, Deep Naryan Mishra, Sujal Reddy Alugubelli +4
Layer-aligned distillation and convergence-based early exit represent two predominant computational efficiency paradigms for transformer inference; yet we establish that they exhib…
cs.LG2026
SAGE: Sign-Adaptive Gradient for Memory-Efficient LLM Optimization
Wooin Lee, Hyun-Tae Kim
The AdamW optimizer, while standard for LLM pretraining, is a critical memory bottleneck, consuming optimizer states equivalent to twice the model's size. Although light-state opti…