Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
FastTPS: An Optimized Method for LLM Token Phase for AI accelerators
Wenzong Yang, Danyang Zhang, Kun Cao +17
The popularity of large language models (LLMs) escalates an ongoing demand for effective inference. However, due to the sequential processing of tokens during the token phase in de…
cs.LG2018
Quantizing Convolutional Neural Networks for Low-Power High-Throughput Inference Engines
Sean O. Settle, Manasa Bollavaram, Paolo D'Alberto +6
Deep learning as a means to inferencing has proliferated thanks to its versatility and ability to approach or exceed human-level accuracy. These computational models have seemingly…