1 paper
Jevin Jiang, Ying Chen, Blake A. Hechtman +2
Large Language Model (LLM) deployment is increasingly shifting to cost-efficient accelerators like Google's Tensor Processing Units (TPUs), prioritizing both performance and total…