collaborators

9 papers

cs.SD2026

dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model

Hankun Wang, Bohan Li, Shi Lian +6

Speech editing for content creation requires precise control over both what an edit should do and where it should apply. Free-form natural language provides a flexible interface fo…

cs.SD2026

dots.tts Technical Report

Shi Lian, Changtao Li, Bohan Li +6

We present dotstts, a 2B-parameter continuous autoregressive text-to-speech (TTS) foundation model that models speech in a continuous latent space. Compared with existing contin…

cs.SD2026

HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding

Bohan Li, Shi Lian, Hankun Wang +6

Unified speech foundation models require a holistic tokenization space that is both learnable by language models and decodable into high-quality waveforms. Existing speech tokenize…

cs.CV2026

Autoregressive Visual Generation Needs a Prologue

Bowen Zheng, Weijian Luo, Guang Yang +2

In this work, we propose Prologue, an approach to bridging the reconstruction-generation gap in autoregressive (AR) image generation. Instead of modifying visual tokens to satisfy…

cs.CV2026

Taming the Entropy Cliff: Variable Codebook Size Quantization for Autoregressive Visual Generation

Bowen Zheng, Weijian Luo, Guang Yang +2

Most discrete visual tokenizers rely on a default design: every position in the sequence shares the same codebook. Researchers try to scale the codebook size to get better reco…

cs.CV2026

Multimodal OCR: Parse Anything from Documents

Handong Zheng, Yumeng Li, Kaile Zhang +22

We present Multimodal OCR (MOCR), a document parsing paradigm that jointly parses text and graphics into unified textual representations. Unlike conventional OCR systems that focus…