papers

Publications (67)

cs.CV2023

Context Autoencoder for Self-Supervised Representation Learning

Xiaokang Chen, Mingyu Ding, Xiaodi Wang +7

We present a novel masked image modeling (MIM) approach, context autoencoder (CAE), for self-supervised representation pretraining. We pretrain an encoder by making predictions in…

cs.CV2023

CAVL: Learning Contrastive and Adaptive Representations of Vision and Language

Shentong Mo, Jingfei Xia, Ihor Markevych

Visual and linguistic pre-training aims to learn vision and language representations together, which can be transferred to visual-linguistic downstream tasks. However, there exists…

cs.CV2023

Audio-Visual Grouping Network for Sound Localization from Mixtures

Shentong Mo, Yapeng Tian

Sound source localization is a typical and challenging task that predicts the location of sound sources in a video. Previous single-source methods mainly used the audio-visual asso…

cs.LG2024

DailyMAE: Towards Pretraining Masked Autoencoders in One Day

Jiantao Wu, Shentong Mo, Sara Atito +3

Recently, masked image modeling (MIM), an important self-supervised learning (SSL) method, has drawn attention for its effectiveness in learning data representation from unlabeled…

cs.CV2024

A Large-scale Medical Visual Task Adaptation Benchmark

Shentong Mo, Xufang Luo, Yansen Wang +1

Visual task adaptation has been demonstrated to be effective in adapting pre-trained Vision Transformers (ViTs) to general downstream visual tasks using specialized learnable layer…

cs.CV2023

Audio-Visual Class-Incremental Learning

Weiguo Pian, Shentong Mo, Yunhui Guo +1

In this paper, we introduce audio-visual class-incremental learning, a class-incremental learning scenario for audio-visual video recognition. We demonstrate that joint audio-visua…

cs.CV2024

Audio-Synchronized Visual Animation

Lin Zhang, Shentong Mo, Yijing Zhang +1

Current visual generation methods can produce high quality videos guided by texts. However, effectively controlling object dynamics remains a challenge. This work explores audio as…

cs.CV2022

Unitail: Detecting, Reading, and Matching in Retail Scene

Fangyi Chen, Han Zhang, Zaiwang Li +7

To make full use of computer vision technology in stores, it is required to consider the actual needs that fit the characteristics of the retail scene. Pursuing this goal, we intro…

cs.CV2026

LVRPO: Language-Visual Alignment with GRPO for Multimodal Understanding and Generation

Shentong Mo, Sukmin Yun

Unified multimodal pretraining has emerged as a promising paradigm for jointly modeling language and vision within a single foundation model. However, existing approaches largely r…

cs.CV2024

Multi-scale Multi-instance Visual Sound Localization and Segmentation

Shentong Mo, Haofan Wang

Visual sound localization is a typical and challenging problem that predicts the location of objects corresponding to the sound source in a video. Previous methods mainly used the…

cs.CV2024

LSPT: Long-term Spatial Prompt Tuning for Visual Representation Learning

Shentong Mo, Yansen Wang, Xufang Luo +1

Visual Prompt Tuning (VPT) techniques have gained prominence for their capacity to adapt pre-trained Vision Transformers (ViTs) to downstream visual tasks using specialized learnab…

cs.CV2026

GMS-CAVP: Improving Audio-Video Correspondence with Multi-Scale Contrastive and Generative Pretraining

Shentong Mo, Zehua Chen, Jun Zhu

Recent advances in video-audio (V-A) understanding and generation have increasingly relied on joint V-A embeddings, which serve as the foundation for tasks such as cross-modal retr…

cs.CV2024

Connecting Joint-Embedding Predictive Architecture with Contrastive Self-supervised Learning

Shentong Mo, Shengbang Tong

In recent advancements in unsupervised visual representation learning, the Joint-Embedding Predictive Architecture (JEPA) has emerged as a significant method for extracting visual…

q-bio.BM2023

Linker-Tuning: Optimizing Continuous Prompts for Heterodimeric Protein Prediction

Shuxian Zou, Hui Li, Shentong Mo +3

Predicting the structure of interacting chains is crucial for understanding biological systems and developing new drugs. Large-scale pre-trained Protein Language Models (PLMs), suc…

cs.CV2024

DMT-JEPA: Discriminative Masked Targets for Joint-Embedding Predictive Architecture

Shentong Mo, Sukmin Yun

The joint-embedding predictive architecture (JEPA) recently has shown impressive results in extracting visual representations from unlabeled imagery under a masking strategy. Howev…

cs.CV2023

Unveiling the Power of Audio-Visual Early Fusion Transformers with Dense Interactions through Masked Modeling

Shentong Mo, Pedro Morgado

Humans possess a remarkable ability to integrate auditory and visual information, enabling a deeper understanding of the surrounding environment. This early fusion of audio and vis…

cs.CV2025

DiffGAP: A Lightweight Diffusion Module in Contrastive Space for Bridging Cross-Model Gap

Shentong Mo, Zehua Chen, Fan Bao +1

Recent works in cross-modal understanding and generation, notably through models like CLAP (Contrastive Language-Audio Pretraining) and CAVP (Contrastive Audio-Visual Pretraining),…

cs.CV2023

Class-Incremental Grouping Network for Continual Audio-Visual Learning

Shentong Mo, Weiguo Pian, Yapeng Tian

Continual learning is a challenging problem in which models need to be trained on non-stationary data across sequential tasks for class-incremental learning. While previous methods…

cs.CV2020

Towards Improving Spatiotemporal Action Recognition in Videos

Shentong Mo, Xiaoqing Tan, Jingfei Xia +1

Spatiotemporal action recognition deals with locating and classifying actions in videos. Motivated by the latest state-of-the-art real-time object detector You Only Watch Once (YOW…

cs.CV2025

Rethinking Positive Pairs in Contrastive Learning

Jiantao Wu, Sara Atito, Zhenhua Feng +3

The training methods in AI do involve semantically distinct pairs of samples. However, their role typically is to enhance the between class separability. The actual notion of simil…

cs.CV2024

Unified Video-Language Pre-training with Synchronized Audio

Shentong Mo, Haofan Wang, Huaxia Li +1

Video-language pre-training is a typical and challenging problem that aims at learning visual and textual representations from large-scale data in a self-supervised way. Existing p…

cs.CV2024

Audio-visual Generalized Zero-shot Learning the Easy Way

Shentong Mo, Pedro Morgado

Audio-visual generalized zero-shot learning is a rapidly advancing domain that seeks to understand the intricate relations between audio and visual cues within videos. The overarch…

cs.LG2026

Modality-Inconsistent Continual Learning of Multimodal Large Language Models

Weiguo Pian, Shijian Deng, Shentong Mo +3

In this paper, we introduce Modality-Inconsistent Continual Learning (MICL), a new continual learning scenario for Multimodal Large Language Models (MLLMs) that involves tasks with…

cs.LG2026

Be Your Own Teacher: Steering Protein Language Models via Unsupervised Reward Optimization

Lanqing Li, Shentong Mo, Yang Yu +1

Protein language models (PLMs) have emerged as powerful tools for controllable biomolecular design, yet their post-training adaptation typically relies on costly wet-lab validation…

cs.CV2022

Multi-Scale Self-Contrastive Learning with Hard Negative Mining for Weakly-Supervised Query-based Video Grounding

Shentong Mo, Daizong Liu, Wei Hu

Query-based video grounding is an important yet challenging task in video understanding, which aims to localize the target segment in an untrimmed video according to a sentence que…

cs.LG2020

Automatic Speech Verification Spoofing Detection

Shentong Mo, Haofan Wang, Pinxu Ren +1

Automatic speech verification (ASV) is the technology to determine the identity of a person based on their voice. While being convenient for identity verification, we should aim fo…

cs.LG2023

High-Modality Multimodal Transformer: Quantifying Modality & Interaction Heterogeneity for High-Modality Representation Learning

Paul Pu Liang, Yiwei Lyu, Xiang Fan +6

Many real-world problems are inherently multimodal, from spoken language, gestures, and paralinguistics humans use to communicate, to force, proprioception, and visual sensors on r…

cs.CV2025

Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm

Lin Zhang, Zefan Cai, Yufan Zhou +10

Recent advances in audio-synchronized visual animation enable control of video content using audios from specific classes. However, existing methods rely heavily on expensive manua…

cs.CV2023

DiT-3D: Exploring Plain Diffusion Transformers for 3D Shape Generation

Shentong Mo, Enze Xie, Ruihang Chu +4

Recent Diffusion Transformers (e.g., DiT) have demonstrated their powerful effectiveness in generating high-quality 2D images. However, it is still being determined whether the Tra…

cs.CV2023

DiffAVA: Personalized Text-to-Audio Generation with Visual Alignment

Shentong Mo, Jing Shi, Yapeng Tian

Text-to-audio (TTA) generation is a recent popular problem that aims to synthesize general audio given text descriptions. Previous methods utilized latent diffusion models to learn…

cs.CV2023

Masked Momentum Contrastive Learning for Zero-shot Semantic Understanding

Jiantao Wu, Shentong Mo, Muhammad Awais +3

Self-supervised pretraining (SSP) has emerged as a popular technique in machine learning, enabling the extraction of meaningful feature representations without labelled data. In th…

cs.CV2023

Beyond Accuracy: Statistical Measures and Benchmark for Evaluation of Representation from Self-Supervised Learning

Jiantao Wu, Shentong Mo, Sara Atito +3

Recently, self-supervised metric learning has raised attention for the potential to learn a generic distance function. It overcomes the limitations of conventional supervised one,…

cs.CV2026

Improving Visual Representation Alignment Generation with GRPO

Shentong Mo, Sukmin Yun

Recent diffusion transformers have demonstrated strong image synthesis capabilities but remain inefficient to train due to weak alignment between generative and discriminative repr…

cs.LG2024

MultiMed: Massively Multimodal and Multitask Medical Understanding

Shentong Mo, Paul Pu Liang

Biomedical data is inherently multimodal, consisting of electronic health records, medical imaging, digital pathology, genome sequencing, wearable sensors, and more. The applicatio…

cs.SD2024

Text-to-Audio Generation Synchronized with Videos

Shentong Mo, Jing Shi, Yapeng Tian

In recent times, the focus on text-to-audio (TTA) generation has intensified, as researchers strive to synthesize audio from textual descriptions. However, most existing methods, t…

cs.LG2024

IoT-LM: Large Multisensory Language Models for the Internet of Things

Shentong Mo, Russ Salakhutdinov, Louis-Philippe Morency +1

The Internet of Things (IoT) network integrating billions of smart physical devices embedded with sensors, software, and communication technologies is a critical and rapidly expand…

cs.LG2023

Variantional autoencoder with decremental information bottleneck for disentanglement

Jiantao Wu, Shentong Mo, Xiang Yang +4

One major challenge of disentanglement learning with variational autoencoders is the trade-off between disentanglement and reconstruction fidelity. Previous studies, which increase…

cs.CV2022

Rethinking Prototypical Contrastive Learning through Alignment, Uniformity and Correlation

Shentong Mo, Zhun Sun, Chao Li

Contrastive self-supervised learning (CSL) with a prototypical regularization has been introduced in learning meaningful representations for downstream tasks that require strong se…

cs.CV2022

Object-wise Masked Autoencoders for Fast Pre-training

Jiantao Wu, Shentong Mo

Self-supervised pre-training for images without labels has recently achieved promising performance in image classification. The success of transformer-based methods, ViT and MAE, d…

cs.CV2026

pMoE: Prompting Diverse Experts Together Wins More in Visual Adaptation

Shentong Mo, Xufang Luo, Dongsheng Li

Parameter-efficient fine-tuning has demonstrated promising results across various visual adaptation tasks, such as classification and segmentation. Typically, prompt tuning techniq…

cs.LG2026

SaDiT: Efficient Protein Backbone Design via Latent Structural Tokenization and Diffusion Transformers

Shentong Mo, Lanqing Li

Generative models for de novo protein backbone design have achieved remarkable success in creating novel protein structures. However, these diffusion-based approaches remain comput…

cs.CV2024

Continual Audio-Visual Sound Separation

Weiguo Pian, Yiyang Nan, Shijian Deng +3

In this paper, we introduce a novel continual audio-visual sound separation task, aiming to continuously separate sound sources for new classes while preserving performance on prev…

cs.SD2023

A Unified Audio-Visual Learning Framework for Localization, Separation, and Recognition

Shentong Mo, Pedro Morgado

The ability to accurately recognize, localize and separate sound sources is fundamental to any audio-visual perception task. Historically, these abilities were tackled separately,…

cs.CV2022

Siamese Prototypical Contrastive Learning

Shentong Mo, Zhun Sun, Chao Li

Contrastive Self-supervised Learning (CSL) is a practical solution that learns meaningful visual representations from massive data in an unsupervised approach. The ordinary CSL emb…

cs.CV2024

The Dynamic Duo of Collaborative Masking and Target for Advanced Masked Autoencoder Learning

Shentong Mo

Masked autoencoders (MAE) have recently succeeded in self-supervised vision representation learning. Previous work mainly applied custom-designed (e.g., random, block-wise) masking…

cs.CV2022

Point3D: tracking actions as moving points with 3D CNNs

Shentong Mo, Jingfei Xia, Xiaoqing Tan +1

Spatio-temporal action recognition has been a challenging task that involves detecting where and when actions occur. Current state-of-the-art action detectors are mostly anchor-bas…

cs.RO2024

We Choose to Go to Space: Agent-driven Human and Multi-Robot Collaboration in Microgravity

Miao Xin, Zhongrui You, Zihan Zhang +7

We present SpaceAgents-1, a system for learning human and multi-robot collaboration (HMRC) strategies under microgravity conditions. Future space exploration requires humans to wor…

cs.CV2024

Efficient 3D Shape Generation via Diffusion Mamba with Bidirectional SSMs

Shentong Mo

Recent advancements in sequence modeling have led to the development of the Mamba architecture, noted for its selective state space approach, offering a promising avenue for effici…

cs.CL2023

Tree of Uncertain Thoughts Reasoning for Large Language Models

Shentong Mo, Miao Xin

While the recently introduced Tree of Thoughts (ToT) has heralded advancements in allowing Large Language Models (LLMs) to reason through foresight and backtracking for global deci…

cs.CV2022

Localizing Visual Sounds the Easy Way

Shentong Mo, Pedro Morgado

Unsupervised audio-visual source localization aims at localizing visible sound sources in a video without relying on ground-truth localization for training. Previous works often se…

q-bio.GN2021

Multi-modal Self-supervised Pre-training for Regulatory Genome Across Cell Types

Shentong Mo, Xi Fu, Chenyang Hong +6

In the genome biology research, regulatory genome modeling is an important topic for many regulatory downstream tasks, such as promoter classification, transaction factor binding s…

cs.CV2024

Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation

Shentong Mo, Yapeng Tian

In recent developments, the Mamba architecture, known for its selective state space approach, has shown potential in the efficient modeling of long sequences. However, its applicat…

cs.CV2023

Exploring Data Augmentations on Self-/Semi-/Fully- Supervised Pre-trained Models

Shentong Mo, Zhun Sun, Chao Li

Data augmentation has become a standard component of vision pre-trained models to capture the invariance between augmented views. In practice, augmentation techniques that mask reg…

cs.CV2023

AV-SAM: Segment Anything Model Meets Audio-Visual Localization and Segmentation

Shentong Mo, Yapeng Tian

Segment Anything Model (SAM) has recently shown its powerful effectiveness in visual segmentation tasks. However, there is less exploration concerning how SAM works on audio-visual…

cs.CV2026

GMAIL: Generative Modality Alignment for generated Image Learning

Shentong Mo, Sukmin Yun

Generative models have made it possible to synthesize highly realistic images, potentially providing an abundant data source for training machine learning models. Despite the advan…

cs.SD2022

A Closer Look at Weakly-Supervised Audio-Visual Source Localization

Shentong Mo, Pedro Morgado

Audio-visual source localization is a challenging task that aims to predict the location of visual sound sources in a video. Since collecting ground-truth annotations of sounding o…

cs.LG2024

MultiIoT: Benchmarking Machine Learning for the Internet of Things

Shentong Mo, Louis-Philippe Morency, Russ Salakhutdinov +1

The next generation of machine learning systems must be adept at perceiving and interacting with the physical world through a diverse array of sensory channels. Commonly referred t…

cs.CV2023

DiffComplete: Diffusion-based Generative 3D Shape Completion

Ruihang Chu, Enze Xie, Shentong Mo +4

We introduce a new diffusion-based approach for shape completion on 3D range scans. Compared with prior deterministic and probabilistic methods, we strike a balance between realism…

cs.CV2024

Aligning Audio-Visual Joint Representations with an Agentic Workflow

Shentong Mo, Yibing Song

Visual content and accompanied audio signals naturally formulate a joint representation to improve audio-visual (AV) related applications. While studies develop various AV represen…

cs.CV2024

MA-AVT: Modality Alignment for Parameter-Efficient Audio-Visual Transformers

Tanvir Mahmud, Shentong Mo, Yapeng Tian +1

Recent advances in pre-trained vision transformers have shown promise in parameter-efficient audio-visual learning without audio pre-training. However, few studies have investigate…

cs.CV2023

Weakly-Supervised Audio-Visual Segmentation

Shentong Mo, Bhiksha Raj

Audio-visual segmentation is a challenging task that aims to predict pixel-level masks for sound sources in a video. Previous work applied a comprehensive manually designed archite…

cs.CV2026

Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows

Shentong Mo, Yibing Song

Coordinated audio generation based on video inputs typically requires a strict audio-visual (AV) alignment, where both semantics and rhythmics of the generated audio segments shall…

cs.CV2023

Fast Training of Diffusion Transformer with Extreme Masking for 3D Point Clouds Generation

Shentong Mo, Enze Xie, Yue Wu +3

Diffusion Transformers have recently shown remarkable effectiveness in generating high-quality 3D point clouds. However, training voxel-based diffusion models for high-resolution 3…

cs.SD2024

Semantic Grouping Network for Audio Source Separation

Shentong Mo, Yapeng Tian

Recently, audio-visual separation approaches have taken advantage of the natural synchronization between the two modalities to boost audio source separation performance. They extra…

cs.LG2021

Learning by Examples Based on Multi-level Optimization

Shentong Mo, Pengtao Xie

Learning by examples, which learns to solve a new problem by looking into how similar problems are solved, is an effective learning method in human learning. When a student learns…

cs.LG2026

APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems

Shentong Mo, Yatao Bian

The paper introduces Atomic Policy Optimization (APO), an unsupervised method that learns to predict 3D structures of atomic systems by optimizing a policy with dual rewards for st…

#unsupervised learning#3d structure prediction#atomic systems#policy optimization
cs.LG2026

Chain of Uncertain Rewards with Large Language Models for Reinforcement Learning

Shentong Mo

Designing effective reward functions is a cornerstone of reinforcement learning (RL), yet it remains a challenging and labor-intensive process due to the inefficiencies and inconsi…