7 citations · 7 across the 2 of their papers we have counts for
2 papers
cs.AR2023★ 7 cited
S: Increasing GPU Utilization during Generative Inference for Higher Throughput
Yunho Jin, Chun-Feng Wu, David Brooks +1
Generating texts with a large language model (LLM) consumes massive amounts of memory. Apart from the already-large model parameters, the key/value (KV) cache that holds informatio…
cs.LG2022
SpeedLimit: Neural Architecture Search for Quantized Transformer Models
Yuji Chai, Luke Bailey, Yunho Jin +5
While research in the field of transformer models has primarily focused on enhancing performance metrics such as accuracy and perplexity, practical applications in industry often n…