2 papers
cs.DC2026
FLYING SERVING: On-the-Fly Parallelism Switching for Large Language Model Serving
Shouwei Gao, Junqi Yin, Feiyi Wang +1
Production LLM serving must simultaneously deliver high throughput, low latency, and sufficient context capacity under non-stationary traffic and mixed request requirements. Data p…
cs.LG2026
LUMOS: Democratizing SciML Workflows with L0-Regularized Learning for Unified Feature and Parameter Adaptation
Shouwei Gao, Xu Zheng, Dongsheng Luo +2
The rapid growth of scientific machine learning (SciML) has accelerated discovery across diverse domains, yet designing effective SciML models remains a challenging task. In practi…