activity
20232026
collaborators

7 papers

cs.CL2026

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis

Wasim Madha, Nityanand Mathur, Hamees Sayed +4

Current text-to-speech systems face a trade-off: autoregres- sive codec language models produce highly intelligible speech but require large-scale models and training data and deco…

cs.AI2026

How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech

Nityanand Mathur, Hamees Sayed, Wasim Madha +4

Style-captioned text-to-speech systems use natural language to control voice characteristics, but how individual words influence acoustic output remains unclear. Understanding this…

cs.AI2026

FlowEdit: Associative Memory for Lifelong Pronunciation Adaptation in Flow-Matching TTS

Harshit Singh, Ayush Pratap Singh, Nityanand Mathur

Flow-matching text-to-speech systems achieve remarkable zero-shot quality but remain static after deployment: pronunciation errors on out-of-vocabulary proper nouns persist unless…

cs.SD2026

SonoEdit: Null-Space Constrained Knowledge Editing for Pronunciation Correction in LLM-Based TTS

Ayush Pratap Singh, Harshit Singh, Nityanand Mathur +2

Neural text-to-speech (TTS) systems systematically mispronounce low-resource proper nouns, particularly non-English names, brands, and geographic locations, due to their underrepre…

cs.CV2024

CraftSVG: Multi-Object Text-to-SVG Synthesis via Layout Guided Diffusion

Ayan Banerjee, Nityanand Mathur, Josep Llados +2

Generating VectorArt from text prompts is a challenging vision task, requiring diverse yet realistic depictions of the seen as well as unseen entities. However, existing research h…

cs.CV2024

DiffuseKronA: A Parameter Efficient Fine-tuning Method for Personalized Diffusion Models

Shyam Marjit, Harshit Singh, Nityanand Mathur +3

In the realm of subject-driven text-to-image (T2I) generative models, recent developments like DreamBooth and BLIP-Diffusion have led to impressive results yet encounter limitation…