9 citations · 11 across the 3 of their papers we have counts for
3 papers
cs.MM2025
3MDiT: Unified Tri-Modal Diffusion Transformer for Text-Driven Synchronized Audio-Video Generation
Yaoru Li, Heyu Si, Federico Landi +8
Text-to-video (T2V) diffusion models have recently achieved impressive visual quality, yet most systems still generate silent clips and treat audio as a secondary concern. Existing…
eess.AS2021★ 2 cited
Location, Location: Enhancing the Evaluation of Text-to-Speech Synthesis Using the Rapid Prosody Transcription Paradigm
Elijah Gutierrez, Pilar Oplustil-Gallegos, Catherine Lai
Text-to-Speech synthesis systems are generally evaluated using Mean Opinion Score (MOS) tests, where listeners score samples of synthetic speech on a Likert scale. A major drawback…
cs.CL2020★ 9 cited
Using previous acoustic context to improve Text-to-Speech synthesis
Pilar Oplustil-Gallegos, Simon King
Many speech synthesis datasets, especially those derived from audiobooks, naturally comprise sequences of utterances. Nevertheless, such data are commonly treated as individual, un…