2 papers
cs.CV2025
Text2Stereo: Repurposing Stable Diffusion for Stereo Generation with Consistency Rewards
Aakash Garg, Libing Zeng, Andrii Tsarov +1
In this paper, we propose a novel diffusion-based approach to generate stereo images given a text prompt. Since stereo image datasets with large baselines are scarce, training a di…
cs.CL2024
Humane Speech Synthesis through Zero-Shot Emotion and Disfluency Generation
Rohan Chaudhury, Mihir Godbole, Aakash Garg +1
Contemporary conversational systems often present a significant limitation: their responses lack the emotional depth and disfluent characteristic of human interactions. This absenc…