8 papers
RefDecoder: Enhancing Visual Generation with Conditional Video Decoding
Xiang Fan, Yuheng Wang, Bohan Fang +2
Video generation powers a vast array of downstream applications. However, while the de facto standard, i.e., latent diffusion models, typically employ heavily conditioned denoising…
Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping
Haoyuan Sun, Jing Wang, Yuxin Song +9
Recently, post-training methods based on reinforcement learning, with a particular focus on Group Relative Policy Optimization (GRPO), have emerged as the robust paradigm for furth…
SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Editing
Xinyao Zhang, Wenkai Dong, Yuxin Song +10
Current instruction-guided video editing models struggle to simultaneously balance precise semantic modifications with faithful motion preservation. While existing approaches rely…
M2P: Improving Visual Foundation Models with Mask-to-Point Weakly-Supervised Learning for Dense Point Tracking
Qiangqiang Wu, Tianyu Yang, Bo Fang +4
Tracking Any Point (TAP) has emerged as a fundamental tool for video understanding. Current approaches adapt Vision Foundation Models (VFMs) like DINOv2 via offline finetuning or t…
ViSS-R1: Self-Supervised Reinforcement Video Reasoning
Bo Fang, Yuxin Song, Qiangqiang Wu +3
Complex video reasoning remains a significant challenge for Multimodal Large Language Models (MLLMs), as current R1-based methodologies often prioritize text-centric reasoning deri…
A Survey on Agentic Multimodal Large Language Models
Huanjin Yao, Ruifei Zhang, Jiaxing Huang +8
With the recent emergence of revolutionary autonomous agentic systems, research community is witnessing a significant shift from traditional static, passive, and domain-specific AI…