4 papers
Avatar V: Scaling Video-Reference Avatar Video Generation
Benjamin Liang, Ce Chen, Desmond Lin +20
Generating avatar videos that are not merely visually similar to a target individual but behaviorally recognizable, faithfully reproducing their talking rhythm, gestural tendencies…
HyGen: Efficient LLM Serving via Elastic Online-Offline Request Co-location
Ting Sun, Penghan Wang, Fan Lai
Large language models (LLMs) have facilitated a wide range of applications with distinct service-level objectives (SLOs), from latency-sensitive online tasks like interactive chatb…
DiSCo: Device-Server Collaborative LLM-Based Text Streaming Services
Ting Sun, Penghan Wang, Fan Lai
The rapid rise of large language models (LLMs) in text streaming services has introduced significant cost and Quality of Experience (QoE) challenges in serving millions of daily re…
Fine-Grained Embedding Dimension Optimization During Training for Recommender Systems
Qinyi Luo, Penghan Wang, Wei Zhang +8
Huge embedding tables in modern deep learning recommender models (DLRM) require prohibitively large memory during training and inference. This paper proposes FIITED, a system to au…