3 papers
cs.CV2022
Learning Audio-Visual embedding for Person Verification in the Wild
Peiwen Sun, Shanshan Zhang, Zishan Liu +4
It has already been observed that audio-visual embedding is more robust than uni-modality embedding for person verification. Here, we proposed a novel audio-visual strategy that co…
cs.SD2021
VRM-Phase I VKW system description of long-short video customizable keyword wakeup challenge
Yougen Yuan, Zhiqiang Lv, Shen Huang +1
Keyword wakeup technology has always been a research hotspot in speech processing, but many related works were done on different datasets. We organized a Chinese long-short video k…
cs.CL2018
Learning Acoustic Word Embeddings with Temporal Context for Query-by-Example Speech Search
Yougen Yuan, Cheung-Chi Leung, Lei Xie +3
We propose to learn acoustic word embeddings with temporal context for query-by-example (QbE) speech search. The temporal context includes the leading and trailing word sequences o…