14 citations · 20 across the 2 of their papers we have counts for
2 papers
cs.CV2022★ 6 cited
Learning Audio-Video Modalities from Image Captions
Arsha Nagrani, Paul Hongsuck Seo, Bryan Seybold +4
A major challenge in text-video and text-audio retrieval is the lack of large-scale training data. This is unlike image-captioning, where datasets are in the order of millions of s…
cs.SD2016★ 14 cited
CNN Architectures for Large-Scale Audio Classification
Shawn Hershey, Sourish Chaudhuri, Daniel P. W. Ellis +10
Convolutional Neural Networks (CNNs) have proven very effective in image classification and show promise for audio. We use various CNN architectures to classify the soundtracks of…