1 citations · 1 across the 3 of their papers we have counts for
3 papers
Audio-Visual Neural Syntax Acquisition
Cheng-I Jeff Lai, Freda Shi, Puyuan Peng +10
We study phrase structure induction from visually-grounded speech. The core idea is to first segment the speech waveform into sequences of word segments, and subsequently induce ph…
Leveraging Pretrained Image-text Models for Improving Audio-Visual Learning
Saurabhchand Bhati, Jesús Villalba, Laureano Moro-Velazquez +2
Visually grounded speech systems learn from paired images and their spoken captions. Recently, there have been attempts to utilize the visually grounded models trained from images…
Regularizing Contrastive Predictive Coding for Speech Applications
Saurabhchand Bhati, Jesús Villalba, Piotr Żelasko +2
Self-supervised methods such as Contrastive predictive Coding (CPC) have greatly improved the quality of the unsupervised representations. These representations significantly reduc…