9 papers
When to Match: A Cost-Balancing Principle for Dynamic Markets
Jie Liu, Hailun Zhang, Jiheng Zhang
Platforms in ridesharing, food delivery, and online gaming must decide not only whom to match but when: immediate matching cuts waiting, while delay thickens the market and improve…
Large-Scale LLM Inference with Heterogeneous Workloads: Prefill-Decode Contention and Asymptotically Optimal Control
Ruihan Lin, Zezhen Ding, Zean Han +1
Large Language Models (LLMs) are rapidly becoming critical infrastructure for enterprise applications, driving unprecedented demand for GPU-based inference services. A key operatio…
Direction-Aware Offline-to-Online Learning in Linear Contextual Bandits
Zean Han, Ruihan Lin, Zezhen Ding +1
Many bandit systems are deployed with offline historical data, such as past logs from earlier policies. Using these data can reduce early online exploration when they remain inform…
Staffing under Taylor's Law: A Unifying Framework for Bridging Square-root and Linear Safety Rules
L. Jeff Hong, Weihuan Huang, Jiheng Zhang +1
Staffing rules are an essential management tool in service industries for meeting target service levels. The square-root safety rule, based on the Poisson arrival assumption, has b…
When to Screen, When to Bypass: LLM-Judges in Resource-Scarce AI-Human Workflow
Ruihan Lin, Jiheng Zhang
AI systems can generate outputs at scale, but most outputs require human approval before release. This creates a bottleneck: humans cannot keep pace with AI-generated volume. A nat…
OR-R1: Automating Modeling and Solving of Operations Research Optimization Problem via Test-Time Reinforcement Learning
Zezhen Ding, Zhen Tan, Jiheng Zhang +1
Optimization modeling and solving are fundamental to the application of Operations Research (OR) in real-world decision making, yet the process of translating natural language prob…