2 papers
cs.LG2026
Pretraining large language models with MXFP4 on Native FP4 Hardware
Musa Cim, Sarthak Arora, Poovaiah Palangappa +4
Why does full-pipeline FP4 training of large language models often diverge, even when forward activations and activation gradients remain stable? We address this question through a…
cs.DC2026
Parallelization Strategies for Dense LLM Deployment: Navigating Through Application-Specific Tradeoffs and Bottlenecks
Burak Topcu, Musa Oguzhan Cim, Poovaiah Palangappa +2
Breakthroughs in the generative AI domain have fueled an explosion of large language model (LLM)-powered applications, whose workloads fundamentally consist of sequences of inferen…