4 papers
OpenMic: A Multi-Agent-Based Stand-Up Comedy Generation System
Yuyang Wu, Hanzhong Cao, Jianhao Chen +1
Chinese stand-up comedy generation goes beyond plain text generation, requiring culturally grounded humor, precise timing, stage-performance cues, and implicit multi-step reasoning…
MACE: A Hybrid LLM Serving System with Colocated SLO-aware Continuous Retraining Alignment
Yufei Li, Yu Fu, Yue Dong +1
Large language models (LLMs) deployed on edge servers are increasingly used in latency-sensitive applications such as personalized assistants, recommendation, and content moderatio…
LeMix: Unified Scheduling for LLM Training and Inference on Multi-GPU Systems
Yufei Li, Zexin Li, Yinglun Zhu +1
Modern deployment of large language models (LLMs) frequently involves both inference serving and continuous retraining to stay aligned with evolving data and user feedback. Common…
Dr Genre: Reinforcement Learning from Decoupled LLM Feedback for Generic Text Rewriting
Yufei Li, John Nham, Ganesh Jawahar +7
Generic text rewriting is a prevalent large language model (LLM) application that covers diverse real-world tasks, such as style transfer, fact correction, and email editing. These…