4 papers
Large-Scale LLM Inference with Heterogeneous Workloads: Prefill-Decode Contention and Asymptotically Optimal Control
Ruihan Lin, Zezhen Ding, Zean Han +1
Large Language Models (LLMs) are rapidly becoming critical infrastructure for enterprise applications, driving unprecedented demand for GPU-based inference services. A key operatio…
Direction-Aware Offline-to-Online Learning in Linear Contextual Bandits
Zean Han, Ruihan Lin, Zezhen Ding +1
Many bandit systems are deployed with offline historical data, such as past logs from earlier policies. Using these data can reduce early online exploration when they remain inform…
When to Screen, When to Bypass: LLM-Judges in Resource-Scarce AI-Human Workflow
Ruihan Lin, Jiheng Zhang
AI systems can generate outputs at scale, but most outputs require human approval before release. This creates a bottleneck: humans cannot keep pace with AI-generated volume. A nat…
Make Optimization Once and for All with Fine-grained Guidance
Mingjia Shi, Ruihan Lin, Xuxi Chen +8
Learning to Optimize (L2O) enhances optimization efficiency with integrated neural networks. L2O paradigms achieve great outcomes, e.g., refitting optimizer, generating unseen solu…