4 papers
Parallel Context Compaction for Long-Horizon LLM Agent Serving
Musa Cim, Burak Topcu, Chita Das +1
Long-horizon LLM agents accumulate growing conversation histories that eventually exceed the model's context window. Context compaction via LLM-based summarization keeps the conver…
Pretraining large language models with MXFP4 on Native FP4 Hardware
Musa Cim, Sarthak Arora, Poovaiah Palangappa +4
Why does full-pipeline FP4 training of large language models often diverge, even when forward activations and activation gradients remain stable? We address this question through a…
Diagnosing FP4 inference: a layer-wise and block-wise sensitivity analysis of NVFP4 and MXFP4
Musa Cim, Burak Topcu, Mahmut Taylan Kandemir
Quantization addresses the high resource demand for large language models (LLMs) by alleviating memory pressure and bandwidth congestion and providing significantly scaled compute…
Parallelization Strategies for Dense LLM Deployment: Navigating Through Application-Specific Tradeoffs and Bottlenecks
Burak Topcu, Musa Oguzhan Cim, Poovaiah Palangappa +2
Breakthroughs in the generative AI domain have fueled an explosion of large language model (LLM)-powered applications, whose workloads fundamentally consist of sequences of inferen…