activity
20242026
collaborators

7 papers

cs.LG2026

Chain of Uncertain Rewards with Large Language Models for Reinforcement Learning

Shentong Mo

Designing effective reward functions is a cornerstone of reinforcement learning (RL), yet it remains a challenging and labor-intensive process due to the inefficiencies and inconsi…

cs.CV2026

LVRPO: Language-Visual Alignment with GRPO for Multimodal Understanding and Generation

Shentong Mo, Sukmin Yun

Unified multimodal pretraining has emerged as a promising paradigm for jointly modeling language and vision within a single foundation model. However, existing approaches largely r…

cs.CV2026

Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows

Shentong Mo, Yibing Song

Coordinated audio generation based on video inputs typically requires a strict audio-visual (AV) alignment, where both semantics and rhythmics of the generated audio segments shall…

cs.CV2025

Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm

Lin Zhang, Zefan Cai, Yufan Zhou +10

Recent advances in audio-synchronized visual animation enable control of video content using audios from specific classes. However, existing methods rely heavily on expensive manua…

cs.CV2024

The Dynamic Duo of Collaborative Masking and Target for Advanced Masked Autoencoder Learning

Shentong Mo

Masked autoencoders (MAE) have recently succeeded in self-supervised vision representation learning. Previous work mainly applied custom-designed (e.g., random, block-wise) masking…

cs.CV2024

Aligning Audio-Visual Joint Representations with an Agentic Workflow

Shentong Mo, Yibing Song

Visual content and accompanied audio signals naturally formulate a joint representation to improve audio-visual (AV) related applications. While studies develop various AV represen…