4 papers
Same Question, Different Answer? Measuring and Mitigating Prompt Privilege for Equitable AI Access
Lier Jin, Lan Hu, Binqi Shen +2
Large language models (LLMs) are increasingly integrated into healthcare, education, public services, and everyday decision making. They should provide comparable assistance regard…
DriftGuard: Safety-Aware Multi-Monitor Detection and Selective Adaptation for Evolving Toxicity Moderation
Yuting Xin, Hanyu Cai, Binqi Shen +2
Automated toxicity moderation systems operate in dynamic online environments where harmful behavior evolves through coded language, shifting targets, and strategic adaptation to en…
The Efficiency Frontier: A Unified Framework for Cost-Performance Optimization in LLM Context Management
Binqi Shen, Lier Jin, Hanyu Cai +2
Large language models (LLMs) increasingly rely on long-context processing, but expanding context windows introduces substantial computational and financial costs. Existing context…
TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput
Xiaoxuan Liu, Jongseok Park, Langxiang Hu +10
Large Language Model (LLM) serving systems batch concurrent user requests to achieve efficient serving. However, in real-world deployments, such inter-request parallelism from batc…