6 papers
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis
Wasim Madha, Nityanand Mathur, Hamees Sayed +4
Current text-to-speech systems face a trade-off: autoregres- sive codec language models produce highly intelligible speech but require large-scale models and training data and deco…
How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech
Nityanand Mathur, Hamees Sayed, Wasim Madha +4
Style-captioned text-to-speech systems use natural language to control voice characteristics, but how individual words influence acoustic output remains unclear. Understanding this…
FlowEdit: Associative Memory for Lifelong Pronunciation Adaptation in Flow-Matching TTS
Harshit Singh, Ayush Pratap Singh, Nityanand Mathur
Flow-matching text-to-speech systems achieve remarkable zero-shot quality but remain static after deployment: pronunciation errors on out-of-vocabulary proper nouns persist unless…
SonoEdit: Null-Space Constrained Knowledge Editing for Pronunciation Correction in LLM-Based TTS
Ayush Pratap Singh, Harshit Singh, Nityanand Mathur +2
Neural text-to-speech (TTS) systems systematically mispronounce low-resource proper nouns, particularly non-English names, brands, and geographic locations, due to their underrepre…
CraftSVG: Multi-Object Text-to-SVG Synthesis via Layout Guided Diffusion
Ayan Banerjee, Nityanand Mathur, Josep Llados +2
Generating VectorArt from text prompts is a challenging vision task, requiring diverse yet realistic depictions of the seen as well as unseen entities. However, existing research h…
CLIPDraw++: Text-to-Sketch Synthesis with Simple Primitives
Nityanand Mathur, Shyam Marjit, Abhra Chaudhuri +1
With the goal of understanding the visual concepts that CLIP associates with text prompts, we show that the latent space of CLIP can be visualized solely in terms of linear transfo…