7 papers
A3 : an Analytical Low-Rank Approximation Framework for Attention
Jeffrey T. H. Wong, Cheng Zhang, Xinye Cao +5
Large language models have demonstrated remarkable performance; however, their massive parameter counts make deployment highly expensive. Low-rank approximation offers a promising…
Quantamination: Dynamic Quantization Leaks Your Data Across the Batch
Hanna Foerster, Ilia Shumailov, Cheng Zhang +3
Dynamic quantization emerged as a practical approach to increase the utilization and efficiency of the machine learning serving flow. Unlike static quantization, which applies quan…
The Llama 4 Herd: Architecture, Training, Evaluation, and Deployment Notes
Redacted by arXiv
This document consolidates publicly reported technical details about Metas Llama 4 model family. It summarizes (i) released variants (Scout and Maverick) and the broader herd conte…
Scaling Laws For Mixed Quantization
Zeyu Cao, Boyang Gu, Cheng Zhang +5
Post-training quantization of Large Language Models (LLMs) has proven effective in reducing the memory and computational requirements for inference. In this study, we focus on a st…
Hardware and Software Platform Inference
Cheng Zhang, Hanna Foerster, Robert D. Mullins +2
It is now a common business practice to buy access to large language model (LLM) inference rather than self-host, because of significant upfront hardware infrastructure and energy…
Locking Machine Learning Models into Hardware
Eleanor Clifford, Adhithya Saravanan, Harry Langford +5
Modern machine learning (ML) models are expensive IP and business competitiveness often depends on keeping this IP confidential. This in turn restricts how these models are deploye…