collaborators

6 papers

cs.LG2026

LatentMoE: Toward Optimal Accuracy per FLOP and Parameter in Mixture of Experts

Venmugil Elango, Nidhi Bhatia, Roger Waleffe +13

Mixture of Experts (MoEs) have become a central component of many state-of-the-art open-source and proprietary large language models. Despite their widespread adoption, it remains…

cs.DC2025

Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens

Yanpeng Yu, Haiyue Ma, Krish Agarwal +10

Expert Parallelism (EP) permits Mixture of Experts (MoE) models to scale beyond a single GPU. To address load imbalance across GPUs in EP, existing approaches aim to balance the nu…

cs.DC2025

Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding

Nidhi Bhatia, Ankit More, Ritika Borkar +7

As LLMs scale to multi-million-token KV histories, real-time autoregressive decoding under tight Token-to-Token Latency (TTL) constraints faces growing pressure. Two core bottlenec…

cs.DC2025

Beyond the Buzz: A Pragmatic Take on Inference Disaggregation

Tiyasa Mitra, Ritika Borkar, Nidhi Bhatia +10

As inference scales to multi-node deployments, disaggregation - splitting inference into distinct phases - offers a promising path to improving the throughput-interactivity Pareto…

cs.LG2025

ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration

Mengting Ai, Tianxin Wei, Yifan Chen +7

Mixture-of-Experts (MoE) Transformer, the backbone architecture of multiple phenomenal language models, leverages sparsity by activating only a fraction of model parameters for eac…

cs.CV2025

Post-Training Quantization for 3D Medical Image Segmentation: A Practical Study on Real Inference Engines

Chongyu Qu, Ritchie Zhao, Ye Yu +6

Quantizing deep neural networks ,reducing the precision (bit-width) of their computations, can remarkably decrease memory usage and accelerate processing, making these models more…