2 papers
cs.LG2026
SAGE: Sign-Adaptive Gradient for Memory-Efficient LLM Optimization
Wooin Lee, Hyun-Tae Kim
The AdamW optimizer, while standard for LLM pretraining, is a critical memory bottleneck, consuming optimizer states equivalent to twice the model's size. Although light-state opti…
cs.CL2021
An Evaluation Dataset and Strategy for Building Robust Multi-turn Response Selection Model
Kijong Han, Seojin Lee, Wooin Lee +2
Multi-turn response selection models have recently shown comparable performance to humans in several benchmark datasets. However, in the real environment, these models often have w…