2 papers
cs.CV2026
HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion
Wenshuo Peng, Kaipeng Zhang
Video-to-audio (V2A) generation faces significant challenges in achieving precise temporal synchronization and high perceptual quality due to the complex, ambiguous relationship be…
cs.SD2026
Semi-Supervised Speech Confidence Detection using Pseudo-Labelling and Whisper Embeddings
Adam Wynn, Jingyun Wang, Xiangyu Tan
Understanding speaker confidence is crucial in educational settings, as it can enhance personalised feedback and improve learning outcomes. This study introduces a novel framework…