collaborators

6 papers

cs.CL2026

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis

Wasim Madha, Nityanand Mathur, Hamees Sayed +4

Current text-to-speech systems face a trade-off: autoregres- sive codec language models produce highly intelligible speech but require large-scale models and training data and deco…

cs.AI2026

How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech

Nityanand Mathur, Hamees Sayed, Wasim Madha +4

Style-captioned text-to-speech systems use natural language to control voice characteristics, but how individual words influence acoustic output remains unclear. Understanding this…

cs.AI2026

FlowEdit: Associative Memory for Lifelong Pronunciation Adaptation in Flow-Matching TTS

Harshit Singh, Ayush Pratap Singh, Nityanand Mathur

Flow-matching text-to-speech systems achieve remarkable zero-shot quality but remain static after deployment: pronunciation errors on out-of-vocabulary proper nouns persist unless…

cs.SD2026

SonoEdit: Null-Space Constrained Knowledge Editing for Pronunciation Correction in LLM-Based TTS

Ayush Pratap Singh, Harshit Singh, Nityanand Mathur +2

Neural text-to-speech (TTS) systems systematically mispronounce low-resource proper nouns, particularly non-English names, brands, and geographic locations, due to their underrepre…

cs.CV2025

CraftSVG: Multi-Object Text-to-SVG Synthesis via Layout Guided Diffusion

Ayan Banerjee, Nityanand Mathur, Josep Llados +2

Generating VectorArt from text prompts is a challenging vision task, requiring diverse yet realistic depictions of the seen as well as unseen entities. However, existing research h…

cs.CV2025

CLIPDraw++: Text-to-Sketch Synthesis with Simple Primitives

Nityanand Mathur, Shyam Marjit, Abhra Chaudhuri +1

With the goal of understanding the visual concepts that CLIP associates with text prompts, we show that the latent space of CLIP can be visualized solely in terms of linear transfo…