1 paper
Schwinn Saereesitthipitak, Mohammed Abdulwahhab, Hannah Zhang +6
Large language model (LLM) inference replicas run across tightly coupled GPUs and serve traffic continuously for weeks. Hardware and software failures are therefore inevitable, and…