5 papers
TailSieve: Partial-Rollout-Guided Tail Routing for LLM Rollouts
Tianqi Xu, Lu Lv, Haoyang Huang +15
Large-scale rollouts have become a core component of modern LLM systems, spanning reinforcement learning (RL) post-training, on-policy distillation (OPD), and sampling-heavy evalua…
OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs
Haoyang Huang, Wenjie Huang, Tianqi Xu +14
Emerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but their long audio-video token sequences introduce substantial memory…
High-Throughput LLM inference on Heterogeneous Clusters
Yi Xiong, Jinqi Huang, Wenjie Huang +6
Nowadays, many companies possess various types of AI accelerators, forming heterogeneous clusters. Efficiently leveraging these clusters for high-throughput large language model (L…
SLO-Aware Scheduling for Large Language Model Inferences
Jinqi Huang, Yi Xiong, Xuebing Yu +4
Large language models (LLMs) have revolutionized applications such as code completion, chatbots, and online classification. To elevate user experiences, service level objectives (S…
WindVE: Collaborative CPU-NPU Vector Embedding
Jinqi Huang, Xuebing Yu, Yi Xiong +4
Retrieval-Augmented Generation is a technology that enhances large language models by integrating information retrieval. In the industry, inference services based on LLMs are highl…