2 papers
cs.LG2026
ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing
Kaiwen Chen, Xin Tan, Minchen Yu +2
Large Reasoning Models (LRMs) are becoming integral to many AI inference systems, enhancing their capabilities with advanced reasoning. However, deploying these models in productio…
cs.LG2024
Echo: Simulating Distributed Training At Scale
Yicheng Feng, Yuetao Chen, Kaiwen Chen +7
Simulation offers unique values for both enumeration and extrapolation purposes, and is becoming increasingly important for managing the massive machine learning (ML) clusters and…