activity
20222024
most citedA Closer Look at Weakly-Supervised Audio-Visual Source Localization

21 citations · 59 across the 17 of their papers we have counts for

collaborators

17 papers

cs.CV20241 cited

A Large-scale Medical Visual Task Adaptation Benchmark

Shentong Mo, Xufang Luo, Yansen Wang +1

Visual task adaptation has been demonstrated to be effective in adapting pre-trained Vision Transformers (ViTs) to general downstream visual tasks using specialized learnable layer…

cs.LG20241 cited

DailyMAE: Towards Pretraining Masked Autoencoders in One Day

Jiantao Wu, Shentong Mo, Sara Atito +3

Recently, masked image modeling (MIM), an important self-supervised learning (SSL) method, has drawn attention for its effectiveness in learning data representation from unlabeled…

cs.SD20242 cited

Text-to-Audio Generation Synchronized with Videos

Shentong Mo, Jing Shi, Yapeng Tian

In recent times, the focus on text-to-audio (TTA) generation has intensified, as researchers strive to synthesize audio from textual descriptions. However, most existing methods, t…

cs.CV2024

LSPT: Long-term Spatial Prompt Tuning for Visual Representation Learning

Shentong Mo, Yansen Wang, Xufang Luo +1

Visual Prompt Tuning (VPT) techniques have gained prominence for their capacity to adapt pre-trained Vision Transformers (ViTs) to downstream visual tasks using specialized learnab…

cs.RO20241 cited

We Choose to Go to Space: Agent-driven Human and Multi-Robot Collaboration in Microgravity

Miao Xin, Zhongrui You, Zihan Zhang +7

We present SpaceAgents-1, a system for learning human and multi-robot collaboration (HMRC) strategies under microgravity conditions. Future space exploration requires humans to wor…

cs.CV2023

Exploring Data Augmentations on Self-/Semi-/Fully- Supervised Pre-trained Models

Shentong Mo, Zhun Sun, Chao Li

Data augmentation has become a standard component of vision pre-trained models to capture the invariance between augmented views. In practice, augmentation techniques that mask reg…