9 papers
dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model
Hankun Wang, Bohan Li, Shi Lian +6
Speech editing for content creation requires precise control over both what an edit should do and where it should apply. Free-form natural language provides a flexible interface fo…
dots.tts Technical Report
Shi Lian, Changtao Li, Bohan Li +6
We present dotstts, a 2B-parameter continuous autoregressive text-to-speech (TTS) foundation model that models speech in a continuous latent space. Compared with existing contin…
HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding
Bohan Li, Shi Lian, Hankun Wang +6
Unified speech foundation models require a holistic tokenization space that is both learnable by language models and decodable into high-quality waveforms. Existing speech tokenize…
Autoregressive Visual Generation Needs a Prologue
Bowen Zheng, Weijian Luo, Guang Yang +2
In this work, we propose Prologue, an approach to bridging the reconstruction-generation gap in autoregressive (AR) image generation. Instead of modifying visual tokens to satisfy…
Taming the Entropy Cliff: Variable Codebook Size Quantization for Autoregressive Visual Generation
Bowen Zheng, Weijian Luo, Guang Yang +2
Most discrete visual tokenizers rely on a default design: every position in the sequence shares the same codebook. Researchers try to scale the codebook size to get better reco…
Multimodal OCR: Parse Anything from Documents
Handong Zheng, Yumeng Li, Kaile Zhang +22
We present Multimodal OCR (MOCR), a document parsing paradigm that jointly parses text and graphics into unified textual representations. Unlike conventional OCR systems that focus…