speech synthesis

AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling

arXiv:2607.12706

summary

AutoSIFT is a text-to-speech framework that separates speaking style into explicit categories (e.g., emotion, age) and residual prosodic details, allowing users to edit specific style attributes while preserving other nuances.

Abstract

State-of-the-art text-to-speech (TTS) models achieve impressive naturalness and expressiveness, yet fine-grained, disentangled control over speaking styles remains challenging. In professional scenarios such as film dubbing, game voice acting, and video content generation, users often need to modify a specific style category, such as emotion, age, or gender, while preserving all others. Existing style-controllable TTS methods typically rely on either text-described styles or speech-reference style transfer, making it difficult to jointly control explicit semantic attributes and preserve subtle, text-undescribed prosodic details. We propose AutoSIFT, a controllable speech generation framework for category-level style editing. AutoSIFT decomposes speaking style into known text-describable categories and unknown residual styles that capture non-verbal prosody and speaker-specific nuances. It consists of a generalized Style Disentangler, which extracts category-aware style prototypes from reference speech, and an Arbitrary Style Infiller, which selectively infills unspecified style categories from the reference. By replacing only text-specified style categories while preserving residual speech-derived styles, AutoSIFT enables natural, expressive, and highly customizable speech generation.

Topics & keywords

#style control#text-to-speech#prosody#voice editing#style disentanglementstyle disentanglerarbitrary style infillercategory-level style editingresidual styleprosodic features