7 papers
Chain of Uncertain Rewards with Large Language Models for Reinforcement Learning
Shentong Mo
Designing effective reward functions is a cornerstone of reinforcement learning (RL), yet it remains a challenging and labor-intensive process due to the inefficiencies and inconsi…
LVRPO: Language-Visual Alignment with GRPO for Multimodal Understanding and Generation
Shentong Mo, Sukmin Yun
Unified multimodal pretraining has emerged as a promising paradigm for jointly modeling language and vision within a single foundation model. However, existing approaches largely r…
Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows
Shentong Mo, Yibing Song
Coordinated audio generation based on video inputs typically requires a strict audio-visual (AV) alignment, where both semantics and rhythmics of the generated audio segments shall…
Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm
Lin Zhang, Zefan Cai, Yufan Zhou +10
Recent advances in audio-synchronized visual animation enable control of video content using audios from specific classes. However, existing methods rely heavily on expensive manua…
The Dynamic Duo of Collaborative Masking and Target for Advanced Masked Autoencoder Learning
Shentong Mo
Masked autoencoders (MAE) have recently succeeded in self-supervised vision representation learning. Previous work mainly applied custom-designed (e.g., random, block-wise) masking…
Aligning Audio-Visual Joint Representations with an Agentic Workflow
Shentong Mo, Yibing Song
Visual content and accompanied audio signals naturally formulate a joint representation to improve audio-visual (AV) related applications. While studies develop various AV represen…