activity
20182024
most citedLearning State-Aware Visual Representations from Audible Interactions

14 citations · 20 across the 6 of their papers we have counts for

collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV2024

Audio-visual Generalized Zero-shot Learning the Easy Way

Shentong Mo, Pedro Morgado

Audio-visual generalized zero-shot learning is a rapidly advancing domain that seeks to understand the intricate relations between audio and visual cues within videos. The overarch…

cs.CV2023

Unveiling the Power of Audio-Visual Early Fusion Transformers with Dense Interactions through Masked Modeling

Shentong Mo, Pedro Morgado

Humans possess a remarkable ability to integrate auditory and visual information, enabling a deeper understanding of the surrounding environment. This early fusion of audio and vis…

cs.CV20231 cited

Why Is Prompt Tuning for Vision-Language Models Robust to Noisy Labels?

Cheng-En Wu, Yu Tian, Haichao Yu +4

Vision-language models such as CLIP learn a generic text-image embedding from large-scale training data. A vision-language model can be adapted to a new classification task through…

cs.CV202214 cited

Learning State-Aware Visual Representations from Audible Interactions

Himangi Mittal, Pedro Morgado, Unnat Jain +1

We propose a self-supervised algorithm to learn representations from egocentric video data. Recently, significant efforts have been made to capture humans interacting with their ow…

cs.CV2022

Localizing Visual Sounds the Easy Way

Shentong Mo, Pedro Morgado

Unsupervised audio-visual source localization aims at localizing visible sound sources in a video without relying on ground-truth localization for training. Previous works often se…

cs.CV20221 cited

The Challenges of Continuous Self-Supervised Learning

Senthil Purushwalkam, Pedro Morgado, Abhinav Gupta

Self-supervised learning (SSL) aims to eliminate one of the major bottlenecks in representation learning - the need for human annotations. As a result, SSL holds the promise to lea…