activity
20192025
most citedAnalyzing Speaker Information in Self-Supervised Models to Improve Zero-Resource Speech Processing

1 citations · 1 across the 8 of their papers we have counts for

collaborators
Showing cs.CLShow all

9 papers · 1 filter

cs.CL2025

The mutual exclusivity bias of bilingual visually grounded speech models

Dan Oneata, Leanne Nortje, Yevgen Matusevych +1

Mutual exclusivity (ME) is a strategy where a novel word is associated with a novel object rather than a familiar one, facilitating language learning in children. Recent work has f…

cs.CL2024

Visually Grounded Speech Models for Low-resource Languages and Cognitive Modelling

Leanne Nortje

This dissertation examines visually grounded speech (VGS) models that learn from unlabelled speech paired with images. It focuses on applications for low-resource languages and und…

cs.CL2024

Improved Visually Prompted Keyword Localisation in Real Low-Resource Settings

Leanne Nortje, Dan Oneata, Gabriel Pirlogeanu +1

Given an image query, visually prompted keyword localisation (VPKL) aims to find occurrences of the depicted word in a speech collection. This can be useful when transcriptions are…

cs.CL2024

Visually Grounded Speech Models have a Mutual Exclusivity Bias

Leanne Nortje, Dan Oneaţă, Yevgen Matusevych +1

When children learn new words, they employ constraints such as the mutual exclusivity (ME) bias: a novel word is mapped to a novel object rather than a familiar one. This bias has…

cs.CL2023

Visually grounded few-shot word acquisition with fewer shots

Leanne Nortje, Benjamin van Niekerk, Herman Kamper

We propose a visually grounded speech model that acquires new words and their visual depictions from just a few word-image example pairs. Given a set of test images and a spoken qu…

cs.CL2022

Towards visually prompted keyword localisation for zero-resource spoken languages

Leanne Nortje, Herman Kamper

Imagine being able to show a system a visual depiction of a keyword and finding spoken utterances that contain this keyword from a zero-resource speech corpus. We formalise this ta…