1 paper
Haoran Jin, Jirong Yang, Zhiheng Zhang +5
Modern ML workloads, with stringent latency and energy constraints, are increasingly hard to run efficiently on homogeneous commodity hardware. We argue that operator-level disaggr…