Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
any4: Learned 4-bit Numeric Representation for LLMs
Mostafa Elhoushi, Jeff Johnson
We present any4, a learned 4-bit weight quantization solution for large language models (LLMs) providing arbitrary numeric representations without requiring pre-processing of weigh…
cs.LG2024
Is Flash Attention Stable?
Alicia Golden, Samuel Hsia, Fei Sun +8
Training large-scale machine learning models poses distinct system challenges, given both the size and complexity of today's workloads. Recently, many organizations training state-…