activity
20192024
most citedAnalyzing Speaker Information in Self-Supervised Models to Improve Zero-Resource Speech Processing

1 citations · 1 across the 6 of their papers we have counts for

collaborators

7 papers

cs.CL2024

Visually Grounded Speech Models for Low-resource Languages and Cognitive Modelling

Leanne Nortje

This dissertation examines visually grounded speech (VGS) models that learn from unlabelled speech paired with images. It focuses on applications for low-resource languages and und…

cs.CL2022

Towards visually prompted keyword localisation for zero-resource spoken languages

Leanne Nortje, Herman Kamper

Imagine being able to show a system a visual depiction of a keyword and finding spoken utterances that contain this keyword from a zero-resource speech corpus. We formalise this ta…

eess.AS20211 cited

Analyzing Speaker Information in Self-Supervised Models to Improve Zero-Resource Speech Processing

Benjamin van Niekerk, Leanne Nortje, Matthew Baas +1

Contrastive predictive coding (CPC) aims to learn representations of speech by distinguishing future observations from a set of negative examples. Previous work has shown that line…

cs.CL2020

Direct multimodal few-shot learning of speech and images

Leanne Nortje, Herman Kamper

We propose direct multimodal few-shot models that learn a shared embedding space of spoken words and images from only a few paired examples. Imagine an agent is shown an image alon…

cs.CL2020

Unsupervised vs. transfer learning for multimodal one-shot matching of speech and images

Leanne Nortje, Herman Kamper

We consider the task of multimodal one-shot speech-image matching. An agent is shown a picture along with a spoken word describing the object in the picture, e.g. cookie, broccoli…

eess.AS2020

Vector-quantized neural networks for acoustic unit discovery in the ZeroSpeech 2020 challenge

Benjamin van Niekerk, Leanne Nortje, Herman Kamper

In this paper, we explore vector quantization for acoustic unit discovery. Leveraging unlabelled data, we aim to learn discrete representations of speech that separate phonetic con…