9 papers
UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods
Yipeng Liu, Chang Liu, Si Shen +16
The deployment of Mixture-of-Experts (MoE) models on production high-bandwidth superpods, such as NVIDIA's NVL72/576 and Huawei's CloudMatrix384, introduces critical challenges bey…
An Efficient and Privacy-Preserving Architecture for Cross-Institutional Collaborative RAG
Chenxin Mao, Shangyu Liu, Zhenzhe Zheng +3
Retrieval-Augmented Generation (RAG) empowers LLMs with external knowledge, making cross-institutional domain-specific knowledge base integration a highly promising deployment para…
Optimizing Feature Extraction for On-device Model Inference with User Behavior Sequences
Chen Gong, Zhenzhe Zheng, Yiliu Chen +3
Machine learning models are widely integrated into modern mobile apps to analyze user behaviors and deliver personalized services. Ensuring low-latency on-device model execution is…
AnyPro: Preference-Preserving Anycast Optimization based on Strategic AS-Path Prepending
Minyuan Zhou, Yuning Chen, Jiaqi Zheng +11
Operating large-scale anycast networks is challenging because client-to-site mappings often misalign with operator's expectation due to opaque inter-domain routing. We present AnyP…
You Need an Encoder for Native Position-Independent Caching
Shiju Zhao, Junhao Hu, Jiaqi Zheng +1
The Key-Value (KV) cache of Large Language Models (LLMs) is prefix-based, making it highly inefficient for processing contexts retrieved in arbitrary order. Position-Independent Ca…
Bi-Level Decision-Focused Causal Learning for Large-Scale Marketing Optimization: Bridging Observational and Experimental Data
Shuli Zhang, Hao Zhou, Jiaqi Zheng +4
Online Internet platforms require sophisticated marketing strategies to optimize user retention and platform revenue -- a classical resource allocation problem. Traditional solutio…