activity
20222024
most citedHuman-like Linguistic Biases in Neural Speech Models: Phonetic Categorization and Phonotactic Constraints in Wav2Vec2.0

3 citations · 6 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CL20243 cited

Human-like Linguistic Biases in Neural Speech Models: Phonetic Categorization and Phonotactic Constraints in Wav2Vec2.0

Marianne de Heer Kloots, Willem Zuidema

What do deep neural speech models know about phonology? Existing work has examined the encoding of individual linguistic units such as phonemes in these models. Here we investigate…

cs.CL20241 cited

Perception of Phonological Assimilation by Neural Speech Recognition Models

Charlotte Pouw, Marianne de Heer Kloots, Afra Alishahi +1

Human listeners effortlessly compensate for phonological changes during speech perception, often unconsciously inferring the intended sounds. For example, listeners infer the under…

cs.CL2023

Transparency at the Source: Evaluating and Interpreting Language Models With Access to the True Distribution

Jaap Jumelet, Willem Zuidema

We present a setup for training, evaluating and interpreting neural language models, that uses artificial, language-like data. The data is generated using a massive probabilistic g…

cs.CL2023

Identifying and Adapting Transformer-Components Responsible for Gender Bias in an English Language Model

Abhijith Chintam, Rahel Beloch, Willem Zuidema +2

Language models (LMs) exhibit and amplify many types of undesirable biases learned from the training data, including gender bias. However, we lack tools for effectively and efficie…

cs.CL2023

Homophone Disambiguation Reveals Patterns of Context Mixing in Speech Transformers

Hosein Mohebbi, Grzegorz Chrupała, Willem Zuidema +1

Transformers have become a key architecture in speech processing, but our understanding of how they build up representations of acoustic and linguistic structure is limited. In thi…

cs.CL2023

Quantifying Context Mixing in Transformers

Hosein Mohebbi, Willem Zuidema, Grzegorz Chrupała +1

Self-attention weights and their transformed variants have been the main source of information for analyzing token-to-token interactions in Transformer-based models. But despite th…