21 citations · 31 across the 5 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2024★ 3 cited
FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
Haojun Xia, Zhen Zheng, Xiaoxia Wu +10
Six-bit quantization (FP6) can effectively reduce the size of large language models (LLMs) and preserve the model quality consistently across varied applications. However, existing…
cs.LG2020★ 21 cited
Larq Compute Engine: Design, Benchmark, and Deploy State-of-the-Art Binarized Neural Networks
Tom Bannink, Arash Bakhtiari, Adam Hillier +5
We introduce Larq Compute Engine, the world's fastest Binarized Neural Network (BNN) inference engine, and use this framework to investigate several important questions about the e…