2 papers
cs.SD2026
CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation
Xiaosu Su, Zihan Sun, Peilei Jia +1
Voice design from natural language descriptions is emerging as a new task in text-to-speech multimodal generation, aiming to synthesize speech with target timbre and speaking style…
cs.CV2025
GAIS: Frame-Level Gated Audio-Visual Integration with Semantic Variance-Scaled Perturbation for Text-Video Retrieval
Bowen Yang, Yun Cao, Chen He +1
Text-to-video retrieval requires precise alignment between language and temporally rich audio-video signals. However, existing methods often emphasize visual cues while underutiliz…