3 papers
eess.AS2025
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Bowen Zhang, Congchao Guo, Geng Yang +17
We introduce MiniMax-Speech, an autoregressive Transformer-based Text-to-Speech (TTS) model that generates high-quality speech. A key innovation is our learnable speaker encoder, w…
eess.IV2025
3DGR-CT: Sparse-View CT Reconstruction with a 3D Gaussian Representation
Yingtai Li, Xueming Fu, Han Li +3
Sparse-view computed tomography (CT) reduces radiation exposure by acquiring fewer projections, making it a valuable tool in clinical scenarios where low-dose radiation is essentia…
cs.CV2024
LoCo: Locally Constrained Training-Free Layout-to-Image Synthesis
Peiang Zhao, Han Li, Ruiyang Jin +1
Recent text-to-image diffusion models have reached an unprecedented level in generating high-quality images. However, their exclusive reliance on textual prompts often falls short…