2 papers
cs.CV2025
GazeCLIP: Enhancing Gaze Estimation Through Text-Guided Multimodal Learning
Jun Wang, Hao Ruan, Liangjian Wen +2
Visual gaze estimation, with its wide-ranging application scenarios, has garnered increasing attention within the research community. Although existing approaches infer gaze solely…
cs.SD2024
Prior-agnostic Multi-scale Contrastive Text-Audio Pre-training for Parallelized TTS Frontend Modeling
Quanxiu Wang, Hui Huang, Mingjie Wang +3
Over the past decade, a series of unflagging efforts have been dedicated to developing highly expressive and controllable text-to-speech (TTS) systems. In general, the holistic TTS…