2 papers
cs.SD2026
MambaVoiceCloning: Efficient and Expressive Text-to-Speech via State-Space Modeling and Diffusion Control
Sahil Kumar, Namrataben Patel, Honggang Wang +1
MambaVoiceCloning (MVC) asks whether the conditioning path of diffusion-based TTS can be made fully SSM-only at inference, removing all attention and explicit RNN-style recurrence…
cs.LG2025
LUMA-RAG: Lifelong Multimodal Agents with Provably Stable Streaming Alignment
Rohan Wandre, Yash Gajewar, Namrata Patel +1
Retrieval-Augmented Generation (RAG) has emerged as the dominant paradigm for grounding large language model outputs in verifiable evidence. However, as modern AI agents transition…