2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.DC2025
Dynamic Rebatching for Efficient Early-Exit Inference with DREX
Xuting Liu, Daniel Alexander, Siva Kesava Reddy Kakarla +2
Early-Exit (EE) is a Large Language Model (LLM) architecture that accelerates inference by allowing easier tokens to be generated using only a subset of the model's layers. However…
cs.NI2023★ 2 cited
Rethinking Machine Learning Collective Communication as a Multi-Commodity Flow Problem
Behnaz Arzani, Siva Kesava Reddy Kakarla, Miguel Castro +3
We show communication schedulers' recent work proposed for ML collectives does not scale to the increasing problem sizes that arise from training larger models. These works also of…