2 papers
cs.DC2026
DHSched: Stateless Control for Stateful Real-Time Avatar Serving
Xin Wang, Haitong Zhang, Xianghong Li +4
Real-time avatar services run long-lived, GPU-backed sessions that transform a continuous text stream into speech, facial motion, rendered video, and RTC media. Elastic operation r…
eess.AS2026
Learnable Classifier-Free Guidance Null Embeddings for Enhanced Controllable Speech Synthesis
Biel Tura Vecino, Yoach Lacombe, Julian Weber +4
Classifier-free Guidance (CFG) is widely adopted in text-to-speech (TTS) systems to enhance generation quality and conditioning fidelity by interpolating between conditioned and un…