14 citations · 31 across the 11 of their papers we have counts for
10 papers
Weakly-Supervised Audio-Visual Segmentation
Shentong Mo, Bhiksha Raj
Audio-visual segmentation is a challenging task that aims to predict pixel-level masks for sound sources in a video. Previous work applied a comprehensive manually designed archite…
Rethinking Prototypical Contrastive Learning through Alignment, Uniformity and Correlation
Shentong Mo, Zhun Sun, Chao Li
Contrastive self-supervised learning (CSL) with a prototypical regularization has been introduced in learning meaningful representations for downstream tasks that require strong se…
Object-wise Masked Autoencoders for Fast Pre-training
Jiantao Wu, Shentong Mo
Self-supervised pre-training for images without labels has recently achieved promising performance in image classification. The success of transformer-based methods, ViT and MAE, d…
Localizing Visual Sounds the Easy Way
Shentong Mo, Pedro Morgado
Unsupervised audio-visual source localization aims at localizing visible sound sources in a video without relying on ground-truth localization for training. Previous works often se…
Point3D: tracking actions as moving points with 3D CNNs
Shentong Mo, Jingfei Xia, Xiaoqing Tan +1
Spatio-temporal action recognition has been a challenging task that involves detecting where and when actions occur. Current state-of-the-art action detectors are mostly anchor-bas…
Multi-Scale Self-Contrastive Learning with Hard Negative Mining for Weakly-Supervised Query-based Video Grounding
Shentong Mo, Daizong Liu, Wei Hu
Query-based video grounding is an important yet challenging task in video understanding, which aims to localize the target segment in an untrimmed video according to a sentence que…