1 citations · 1 across the 2 of their papers we have counts for
4 papers
Sangam: Efficiently Serving Diffusion LLMs with the AR Stack
Nitin Kedia, Saurabh Agarwal, Myungjin Lee +1
Diffusion language models (dLLMs) generate text by iteratively denoising a masked response and can commit multiple output positions per model invocation. Their bidirectional attent…
Harmonia: End-to-End RAG Serving Optimization
Saurabh Agarwal, Bodun Hu, Luis Pabon +3
Retrieval-Augmented Generation (RAG) improves the reliability of large language models by integrating external knowledge, but serving RAG pipelines efficiently is challenging becau…
FedMOA: Federated GRPO for Personalized Reasoning LLMs under Heterogeneous Rewards
Ziyao Wang, Daeun Jung, Yexiao He +4
Group Relative Policy Optimization (GRPO) has recently emerged as an effective approach for improving the reasoning capabilities of large language models through online multi-objec…
Prada: Black-Box LLM Adaptation with Private Data on Resource-Constrained Devices
Ziyao Wang, Yexiao He, Zheyu Shen +4
In recent years, Large Language Models (LLMs) have demonstrated remarkable abilities in various natural language processing tasks. However, adapting these models to specialized dom…