2 papers
cs.DC2024
APEX: An Extensible and Dynamism-Aware Simulator for Automated Parallel Execution in LLM Serving
Yi-Chien Lin, Woosuk Kwon, Ronald Pineda +1
Efficiently serving Large Language Models (LLMs) requires selecting an optimal parallel execution plan, balancing computation, memory, and communication overhead. However, determin…
cs.LG2020
Efficient Algorithms for Device Placement of DNN Graph Operators
Jakub Tarnawski, Amar Phanishayee, Nikhil R. Devanur +2
Modern machine learning workloads use large models, with complex structures, that are very expensive to execute. The devices that execute complex models are becoming increasingly h…