most citedAASIST: Audio Anti-Spoofing using Integrated Spectro-Temporal Graph Attention Networks

4 citations · 13 across the 10 of their papers we have counts for

collaborators

11 papers

eess.AS2022

Metric Learning for User-defined Keyword Spotting

Jaemin Jung, Youkyum Kim, Jihwan Park +4

The goal of this work is to detect new spoken terms defined by users. While most previous works address Keyword Spotting (KWS) as a closed-set classification problem, this limits t…

cs.SD20224 cited

Large-scale learning of generalised representations for speaker recognition

Jee-weon Jung, Hee-Soo Heo, Bong-Jin Lee +5

The objective of this work is to develop a speaker recognition model to be used in diverse scenarios. We hypothesise that two components should be adequately configured to build su…

cs.SD2022

In search of strong embedding extractors for speaker diarisation

Jee-weon Jung, Hee-Soo Heo, Bong-Jin Lee +5

Speaker embedding extractors (EEs), which map input audio to a speaker discriminant latent space, are of paramount importance in speaker diarisation. However, there are several cha…

cs.SD20221 cited

Baseline Systems for the First Spoofing-Aware Speaker Verification Challenge: Score and Embedding Fusion

Hye-jin Shim, Hemlata Tak, Xuechen Liu +12

Deep learning has brought impressive progress in the study of both automatic speaker verification (ASV) and spoofing countermeasures (CM). Although solutions are mutually dependent…

eess.AS20223 cited

Pushing the limits of raw waveform speaker recognition

Jee-weon Jung, You Jin Kim, Hee-Soo Heo +3

In recent years, speaker recognition systems based on raw waveform inputs have received increasing attention. However, the performance of such systems are typically inferior to the…

eess.AS2021

Multi-scale speaker embedding-based graph attention networks for speaker diarisation

Youngki Kwon, Hee-Soo Heo, Jee-weon Jung +3

The objective of this work is effective speaker diarisation using multi-scale speaker embeddings. Typically, there is a trade-off between the ability to recognise short speaker seg…