4 papers
Scalable Pretraining of Large Mixture of Experts Language Models on Aurora Super Computer
Dharma Teja Vooturi, Dhiraj Kalamkar, Dipankar Das +1
Pretraining Large Language Models (LLMs) from scratch requires massive amount of compute. Aurora super computer is an ExaScale machine with 127,488 Intel PVC (Ponte Vechio) GPU til…
Pushing the Envelope of LLM Inference on AI-PC and Intel GPUs
Evangelos Georganas, Dhiraj Kalamkar, Alexander Heinecke
The advent of ultra-low-bit LLM models (1/1.58/2-bit), which match the perplexity and end-task performance of their full-precision counterparts using the same model size, is usheri…
ML-SpecQD: Multi-Level Speculative Decoding with Quantized Drafts
Evangelos Georganas, Dhiraj Kalamkar, Alexander Kozlov +1
Speculative decoding (SD) has emerged as a method to accelerate LLM inference without sacrificing any accuracy over the 16-bit model inference. In a typical SD setup, the idea is t…
DQRM: Deep Quantized Recommendation Models
Yang Zhou, Zhen Dong, Ellick Chan +3
Large-scale recommendation models are currently the dominant workload for many large Internet companies. These recommenders are characterized by massive embedding tables that are s…