Publications (67)
Context Autoencoder for Self-Supervised Representation Learning
Xiaokang Chen, Mingyu Ding, Xiaodi Wang +7
We present a novel masked image modeling (MIM) approach, context autoencoder (CAE), for self-supervised representation pretraining. We pretrain an encoder by making predictions in…
CAVL: Learning Contrastive and Adaptive Representations of Vision and Language
Shentong Mo, Jingfei Xia, Ihor Markevych
Visual and linguistic pre-training aims to learn vision and language representations together, which can be transferred to visual-linguistic downstream tasks. However, there exists…
Audio-Visual Grouping Network for Sound Localization from Mixtures
Shentong Mo, Yapeng Tian
Sound source localization is a typical and challenging task that predicts the location of sound sources in a video. Previous single-source methods mainly used the audio-visual asso…
DailyMAE: Towards Pretraining Masked Autoencoders in One Day
Jiantao Wu, Shentong Mo, Sara Atito +3
Recently, masked image modeling (MIM), an important self-supervised learning (SSL) method, has drawn attention for its effectiveness in learning data representation from unlabeled…
A Large-scale Medical Visual Task Adaptation Benchmark
Shentong Mo, Xufang Luo, Yansen Wang +1
Visual task adaptation has been demonstrated to be effective in adapting pre-trained Vision Transformers (ViTs) to general downstream visual tasks using specialized learnable layer…
Audio-Visual Class-Incremental Learning
Weiguo Pian, Shentong Mo, Yunhui Guo +1
In this paper, we introduce audio-visual class-incremental learning, a class-incremental learning scenario for audio-visual video recognition. We demonstrate that joint audio-visua…
Audio-Synchronized Visual Animation
Lin Zhang, Shentong Mo, Yijing Zhang +1
Current visual generation methods can produce high quality videos guided by texts. However, effectively controlling object dynamics remains a challenge. This work explores audio as…
Unitail: Detecting, Reading, and Matching in Retail Scene
Fangyi Chen, Han Zhang, Zaiwang Li +7
To make full use of computer vision technology in stores, it is required to consider the actual needs that fit the characteristics of the retail scene. Pursuing this goal, we intro…
LVRPO: Language-Visual Alignment with GRPO for Multimodal Understanding and Generation
Shentong Mo, Sukmin Yun
Unified multimodal pretraining has emerged as a promising paradigm for jointly modeling language and vision within a single foundation model. However, existing approaches largely r…
Multi-scale Multi-instance Visual Sound Localization and Segmentation
Shentong Mo, Haofan Wang
Visual sound localization is a typical and challenging problem that predicts the location of objects corresponding to the sound source in a video. Previous methods mainly used the…
LSPT: Long-term Spatial Prompt Tuning for Visual Representation Learning
Shentong Mo, Yansen Wang, Xufang Luo +1
Visual Prompt Tuning (VPT) techniques have gained prominence for their capacity to adapt pre-trained Vision Transformers (ViTs) to downstream visual tasks using specialized learnab…
GMS-CAVP: Improving Audio-Video Correspondence with Multi-Scale Contrastive and Generative Pretraining
Shentong Mo, Zehua Chen, Jun Zhu
Recent advances in video-audio (V-A) understanding and generation have increasingly relied on joint V-A embeddings, which serve as the foundation for tasks such as cross-modal retr…
Connecting Joint-Embedding Predictive Architecture with Contrastive Self-supervised Learning
Shentong Mo, Shengbang Tong
In recent advancements in unsupervised visual representation learning, the Joint-Embedding Predictive Architecture (JEPA) has emerged as a significant method for extracting visual…
Linker-Tuning: Optimizing Continuous Prompts for Heterodimeric Protein Prediction
Shuxian Zou, Hui Li, Shentong Mo +3
Predicting the structure of interacting chains is crucial for understanding biological systems and developing new drugs. Large-scale pre-trained Protein Language Models (PLMs), suc…
DMT-JEPA: Discriminative Masked Targets for Joint-Embedding Predictive Architecture
Shentong Mo, Sukmin Yun
The joint-embedding predictive architecture (JEPA) recently has shown impressive results in extracting visual representations from unlabeled imagery under a masking strategy. Howev…
Unveiling the Power of Audio-Visual Early Fusion Transformers with Dense Interactions through Masked Modeling
Shentong Mo, Pedro Morgado
Humans possess a remarkable ability to integrate auditory and visual information, enabling a deeper understanding of the surrounding environment. This early fusion of audio and vis…
DiffGAP: A Lightweight Diffusion Module in Contrastive Space for Bridging Cross-Model Gap
Shentong Mo, Zehua Chen, Fan Bao +1
Recent works in cross-modal understanding and generation, notably through models like CLAP (Contrastive Language-Audio Pretraining) and CAVP (Contrastive Audio-Visual Pretraining),…
Class-Incremental Grouping Network for Continual Audio-Visual Learning
Shentong Mo, Weiguo Pian, Yapeng Tian
Continual learning is a challenging problem in which models need to be trained on non-stationary data across sequential tasks for class-incremental learning. While previous methods…
Towards Improving Spatiotemporal Action Recognition in Videos
Shentong Mo, Xiaoqing Tan, Jingfei Xia +1
Spatiotemporal action recognition deals with locating and classifying actions in videos. Motivated by the latest state-of-the-art real-time object detector You Only Watch Once (YOW…
Rethinking Positive Pairs in Contrastive Learning
Jiantao Wu, Sara Atito, Zhenhua Feng +3
The training methods in AI do involve semantically distinct pairs of samples. However, their role typically is to enhance the between class separability. The actual notion of simil…
Unified Video-Language Pre-training with Synchronized Audio
Shentong Mo, Haofan Wang, Huaxia Li +1
Video-language pre-training is a typical and challenging problem that aims at learning visual and textual representations from large-scale data in a self-supervised way. Existing p…
Audio-visual Generalized Zero-shot Learning the Easy Way
Shentong Mo, Pedro Morgado
Audio-visual generalized zero-shot learning is a rapidly advancing domain that seeks to understand the intricate relations between audio and visual cues within videos. The overarch…
Modality-Inconsistent Continual Learning of Multimodal Large Language Models
Weiguo Pian, Shijian Deng, Shentong Mo +3
In this paper, we introduce Modality-Inconsistent Continual Learning (MICL), a new continual learning scenario for Multimodal Large Language Models (MLLMs) that involves tasks with…
Be Your Own Teacher: Steering Protein Language Models via Unsupervised Reward Optimization
Lanqing Li, Shentong Mo, Yang Yu +1
Protein language models (PLMs) have emerged as powerful tools for controllable biomolecular design, yet their post-training adaptation typically relies on costly wet-lab validation…
Multi-Scale Self-Contrastive Learning with Hard Negative Mining for Weakly-Supervised Query-based Video Grounding
Shentong Mo, Daizong Liu, Wei Hu
Query-based video grounding is an important yet challenging task in video understanding, which aims to localize the target segment in an untrimmed video according to a sentence que…
Automatic Speech Verification Spoofing Detection
Shentong Mo, Haofan Wang, Pinxu Ren +1
Automatic speech verification (ASV) is the technology to determine the identity of a person based on their voice. While being convenient for identity verification, we should aim fo…
High-Modality Multimodal Transformer: Quantifying Modality & Interaction Heterogeneity for High-Modality Representation Learning
Paul Pu Liang, Yiwei Lyu, Xiang Fan +6
Many real-world problems are inherently multimodal, from spoken language, gestures, and paralinguistics humans use to communicate, to force, proprioception, and visual sensors on r…
Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm
Lin Zhang, Zefan Cai, Yufan Zhou +10
Recent advances in audio-synchronized visual animation enable control of video content using audios from specific classes. However, existing methods rely heavily on expensive manua…
DiT-3D: Exploring Plain Diffusion Transformers for 3D Shape Generation
Shentong Mo, Enze Xie, Ruihang Chu +4
Recent Diffusion Transformers (e.g., DiT) have demonstrated their powerful effectiveness in generating high-quality 2D images. However, it is still being determined whether the Tra…
DiffAVA: Personalized Text-to-Audio Generation with Visual Alignment
Shentong Mo, Jing Shi, Yapeng Tian
Text-to-audio (TTA) generation is a recent popular problem that aims to synthesize general audio given text descriptions. Previous methods utilized latent diffusion models to learn…
Masked Momentum Contrastive Learning for Zero-shot Semantic Understanding
Jiantao Wu, Shentong Mo, Muhammad Awais +3
Self-supervised pretraining (SSP) has emerged as a popular technique in machine learning, enabling the extraction of meaningful feature representations without labelled data. In th…
Beyond Accuracy: Statistical Measures and Benchmark for Evaluation of Representation from Self-Supervised Learning
Jiantao Wu, Shentong Mo, Sara Atito +3
Recently, self-supervised metric learning has raised attention for the potential to learn a generic distance function. It overcomes the limitations of conventional supervised one,…
Improving Visual Representation Alignment Generation with GRPO
Shentong Mo, Sukmin Yun
Recent diffusion transformers have demonstrated strong image synthesis capabilities but remain inefficient to train due to weak alignment between generative and discriminative repr…
MultiMed: Massively Multimodal and Multitask Medical Understanding
Shentong Mo, Paul Pu Liang
Biomedical data is inherently multimodal, consisting of electronic health records, medical imaging, digital pathology, genome sequencing, wearable sensors, and more. The applicatio…
Text-to-Audio Generation Synchronized with Videos
Shentong Mo, Jing Shi, Yapeng Tian
In recent times, the focus on text-to-audio (TTA) generation has intensified, as researchers strive to synthesize audio from textual descriptions. However, most existing methods, t…
IoT-LM: Large Multisensory Language Models for the Internet of Things
Shentong Mo, Russ Salakhutdinov, Louis-Philippe Morency +1
The Internet of Things (IoT) network integrating billions of smart physical devices embedded with sensors, software, and communication technologies is a critical and rapidly expand…
Variantional autoencoder with decremental information bottleneck for disentanglement
Jiantao Wu, Shentong Mo, Xiang Yang +4
One major challenge of disentanglement learning with variational autoencoders is the trade-off between disentanglement and reconstruction fidelity. Previous studies, which increase…
Rethinking Prototypical Contrastive Learning through Alignment, Uniformity and Correlation
Shentong Mo, Zhun Sun, Chao Li
Contrastive self-supervised learning (CSL) with a prototypical regularization has been introduced in learning meaningful representations for downstream tasks that require strong se…
Object-wise Masked Autoencoders for Fast Pre-training
Jiantao Wu, Shentong Mo
Self-supervised pre-training for images without labels has recently achieved promising performance in image classification. The success of transformer-based methods, ViT and MAE, d…
pMoE: Prompting Diverse Experts Together Wins More in Visual Adaptation
Shentong Mo, Xufang Luo, Dongsheng Li
Parameter-efficient fine-tuning has demonstrated promising results across various visual adaptation tasks, such as classification and segmentation. Typically, prompt tuning techniq…
SaDiT: Efficient Protein Backbone Design via Latent Structural Tokenization and Diffusion Transformers
Shentong Mo, Lanqing Li
Generative models for de novo protein backbone design have achieved remarkable success in creating novel protein structures. However, these diffusion-based approaches remain comput…
Continual Audio-Visual Sound Separation
Weiguo Pian, Yiyang Nan, Shijian Deng +3
In this paper, we introduce a novel continual audio-visual sound separation task, aiming to continuously separate sound sources for new classes while preserving performance on prev…
A Unified Audio-Visual Learning Framework for Localization, Separation, and Recognition
Shentong Mo, Pedro Morgado
The ability to accurately recognize, localize and separate sound sources is fundamental to any audio-visual perception task. Historically, these abilities were tackled separately,…
Siamese Prototypical Contrastive Learning
Shentong Mo, Zhun Sun, Chao Li
Contrastive Self-supervised Learning (CSL) is a practical solution that learns meaningful visual representations from massive data in an unsupervised approach. The ordinary CSL emb…
The Dynamic Duo of Collaborative Masking and Target for Advanced Masked Autoencoder Learning
Shentong Mo
Masked autoencoders (MAE) have recently succeeded in self-supervised vision representation learning. Previous work mainly applied custom-designed (e.g., random, block-wise) masking…
Point3D: tracking actions as moving points with 3D CNNs
Shentong Mo, Jingfei Xia, Xiaoqing Tan +1
Spatio-temporal action recognition has been a challenging task that involves detecting where and when actions occur. Current state-of-the-art action detectors are mostly anchor-bas…
We Choose to Go to Space: Agent-driven Human and Multi-Robot Collaboration in Microgravity
Miao Xin, Zhongrui You, Zihan Zhang +7
We present SpaceAgents-1, a system for learning human and multi-robot collaboration (HMRC) strategies under microgravity conditions. Future space exploration requires humans to wor…
Efficient 3D Shape Generation via Diffusion Mamba with Bidirectional SSMs
Shentong Mo
Recent advancements in sequence modeling have led to the development of the Mamba architecture, noted for its selective state space approach, offering a promising avenue for effici…
Tree of Uncertain Thoughts Reasoning for Large Language Models
Shentong Mo, Miao Xin
While the recently introduced Tree of Thoughts (ToT) has heralded advancements in allowing Large Language Models (LLMs) to reason through foresight and backtracking for global deci…
Localizing Visual Sounds the Easy Way
Shentong Mo, Pedro Morgado
Unsupervised audio-visual source localization aims at localizing visible sound sources in a video without relying on ground-truth localization for training. Previous works often se…
Multi-modal Self-supervised Pre-training for Regulatory Genome Across Cell Types
Shentong Mo, Xi Fu, Chenyang Hong +6
In the genome biology research, regulatory genome modeling is an important topic for many regulatory downstream tasks, such as promoter classification, transaction factor binding s…
Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation
Shentong Mo, Yapeng Tian
In recent developments, the Mamba architecture, known for its selective state space approach, has shown potential in the efficient modeling of long sequences. However, its applicat…
Exploring Data Augmentations on Self-/Semi-/Fully- Supervised Pre-trained Models
Shentong Mo, Zhun Sun, Chao Li
Data augmentation has become a standard component of vision pre-trained models to capture the invariance between augmented views. In practice, augmentation techniques that mask reg…
AV-SAM: Segment Anything Model Meets Audio-Visual Localization and Segmentation
Shentong Mo, Yapeng Tian
Segment Anything Model (SAM) has recently shown its powerful effectiveness in visual segmentation tasks. However, there is less exploration concerning how SAM works on audio-visual…
GMAIL: Generative Modality Alignment for generated Image Learning
Shentong Mo, Sukmin Yun
Generative models have made it possible to synthesize highly realistic images, potentially providing an abundant data source for training machine learning models. Despite the advan…
A Closer Look at Weakly-Supervised Audio-Visual Source Localization
Shentong Mo, Pedro Morgado
Audio-visual source localization is a challenging task that aims to predict the location of visual sound sources in a video. Since collecting ground-truth annotations of sounding o…
MultiIoT: Benchmarking Machine Learning for the Internet of Things
Shentong Mo, Louis-Philippe Morency, Russ Salakhutdinov +1
The next generation of machine learning systems must be adept at perceiving and interacting with the physical world through a diverse array of sensory channels. Commonly referred t…
DiffComplete: Diffusion-based Generative 3D Shape Completion
Ruihang Chu, Enze Xie, Shentong Mo +4
We introduce a new diffusion-based approach for shape completion on 3D range scans. Compared with prior deterministic and probabilistic methods, we strike a balance between realism…
Aligning Audio-Visual Joint Representations with an Agentic Workflow
Shentong Mo, Yibing Song
Visual content and accompanied audio signals naturally formulate a joint representation to improve audio-visual (AV) related applications. While studies develop various AV represen…
MA-AVT: Modality Alignment for Parameter-Efficient Audio-Visual Transformers
Tanvir Mahmud, Shentong Mo, Yapeng Tian +1
Recent advances in pre-trained vision transformers have shown promise in parameter-efficient audio-visual learning without audio pre-training. However, few studies have investigate…
Weakly-Supervised Audio-Visual Segmentation
Shentong Mo, Bhiksha Raj
Audio-visual segmentation is a challenging task that aims to predict pixel-level masks for sound sources in a video. Previous work applied a comprehensive manually designed archite…
Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows
Shentong Mo, Yibing Song
Coordinated audio generation based on video inputs typically requires a strict audio-visual (AV) alignment, where both semantics and rhythmics of the generated audio segments shall…
Fast Training of Diffusion Transformer with Extreme Masking for 3D Point Clouds Generation
Shentong Mo, Enze Xie, Yue Wu +3
Diffusion Transformers have recently shown remarkable effectiveness in generating high-quality 3D point clouds. However, training voxel-based diffusion models for high-resolution 3…
Semantic Grouping Network for Audio Source Separation
Shentong Mo, Yapeng Tian
Recently, audio-visual separation approaches have taken advantage of the natural synchronization between the two modalities to boost audio source separation performance. They extra…
Learning by Examples Based on Multi-level Optimization
Shentong Mo, Pengtao Xie
Learning by examples, which learns to solve a new problem by looking into how similar problems are solved, is an effective learning method in human learning. When a student learns…
APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems
Shentong Mo, Yatao Bian
The paper introduces Atomic Policy Optimization (APO), an unsupervised method that learns to predict 3D structures of atomic systems by optimizing a policy with dual rewards for st…
Chain of Uncertain Rewards with Large Language Models for Reinforcement Learning
Shentong Mo
Designing effective reward functions is a cornerstone of reinforcement learning (RL), yet it remains a challenging and labor-intensive process due to the inefficiencies and inconsi…