2 papers
cs.SD2026
Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers
Dongseong Hwang, Prasanth Yadla, Kaan Elgin +8
Siri Expressive Voices synthesize rich, configurable speech in real time and entirely on device, powered by AFM 3 Core Advanced, Apple's most powerful on-device foundation model. T…
cs.AI2024
Towards Low-bit Communication for Tensor Parallel LLM Inference
Harry Dong, Tyler Johnson, Minsik Cho +1
Tensor parallelism provides an effective way to increase server large language model (LLM) inference efficiency despite adding an additional communication cost. However, as server…