Publications (19)
Wan: Open and Advanced Large-Scale Video Generative Models
Team Wan, Ang Wang, Baole Ai +58
This report presents Wan, a comprehensive and open suite of video foundation models designed to push the boundaries of video generation. Built upon the mainstream diffusion transfo…
Panchromatic and Multispectral Image Fusion via Alternating Reverse Filtering Network
Keyu Yan, Man Zhou, Jie Huang +4
Panchromatic (PAN) and multi-spectral (MS) image fusion, named Pan-sharpening, refers to super-resolve the low-resolution (LR) multi-spectral (MS) images in the spatial domain to g…
Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoE
Haoyou Deng, Keyu Yan, Chaojie Mao +4
Mixture-of-Experts (MoE) architectures have emerged as a powerful paradigm for scaling diffusion models in visual generation. Recent advancements have focused on adaptively allocat…
DenseGRPO: From Sparse to Dense Reward for Flow Matching Model Alignment
Haoyou Deng, Keyu Yan, Chaojie Mao +4
Recent GRPO-based approaches built on flow matching models have shown remarkable improvements in human preference alignment for text-to-image generation. Nevertheless, they still s…
Perception, Verdict, and Evolution: Hindsight-Driven Self-Refining Forensics Agent for AI-Generated Image Detection
Yangjun Wu, Keyu Yan, Yu Liu +5
The rapid advancement of generative models presents a significant challenge to existing deepfake detection methods, particularly given the widespread dissemination of highly realis…
Training-Free Large Model Priors for Multiple-in-One Image Restoration
Xuanhua He, Lang Li, Yingying Wang +7
Image restoration aims to reconstruct the latent clear images from their degraded versions. Despite the notable achievement, existing methods predominantly focus on handling specif…
Pan-Mamba: Effective pan-sharpening with State Space Model
Xuanhua He, Ke Cao, Keyu Yan +4
Pan-sharpening involves integrating information from low-resolution multi-spectral and high-resolution panchromatic images to generate high-resolution multi-spectral counterparts.…
Enhancing RAW-to-sRGB with Decoupled Style Structure in Fourier Domain
Xuanhua He, Tao Hu, Guoli Wang +9
RAW to sRGB mapping, which aims to convert RAW images from smartphones into RGB form equivalent to that of Digital Single-Lens Reflex (DSLR) cameras, has become an important area o…
Wan-Image: Pushing the Boundaries of Generative Visual Intelligence
Chaojie Mao, Chen-Wei Xie, Chongyang Zhong +55
We present Wan-Image, a unified visual generation system explicitly engineered to paradigm-shift image generation models from casual synthesizers into professional-grade productivi…
Wan-Weaver: Interleaved Multi-modal Generation via Decoupled Training
Jinbo Xing, Zeyinzi Jiang, Yuxiang Tuo +15
Recent unified models have made unprecedented progress in both understanding and generation. However, while most of them accept multi-modal inputs, they typically produce only sing…
Cross-database non-frontal facial expression recognition based on transductive deep transfer learning
Keyu Yan, Wenming Zheng, Tong Zhang +2
Cross-database non-frontal expression recognition is a very meaningful but rather difficult subject in the fields of computer vision and affect computing. In this paper, we propose…
Low-resource Language Discrimination Towards Chinese Dialects with Transfer learning and Data Augmentation
Fan Xu, Yangjie Dan, Keyu Yan +2
Chinese dialects discrimination is a challenging natural language processing task due to scarce annotation resource. In this article, we develop a novel Chinese dialects discrimina…
Memory-augmented Deep Unfolding Network for Guided Image Super-resolution
Man Zhou, Keyu Yan, Jinshan Pan +3
Guided image super-resolution (GISR) aims to obtain a high-resolution (HR) target image by enhancing the spatial resolution of a low-resolution (LR) target image under the guidance…
ID-Animator: Zero-Shot Identity-Preserving Human Video Generation
Xuanhua He, Quande Liu, Shengju Qian +5
Generating high-fidelity human video with specified identities has attracted significant attention in the content generation community. However, existing techniques struggle to str…
VIPO: Value Function Inconsistency Penalized Offline Reinforcement Learning
Xuyang Chen, Keyu Yan, Guojian Wang +1
Offline reinforcement learning (RL) learns effective policies from pre-collected datasets, offering a practical solution for applications where online interactions are risky or cos…
One-Step Sampler for Boltzmann Distributions via Drifting
Wenhan Cao, Keyu Yan, Lin Zhao
We present a drifting-based framework for amortized sampling of Boltzmann distributions defined by energy functions. The method trains a one-step neural generator by projecting sam…
BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs
Zhantao Yang, Ruili Feng, Keyu Yan +13
Advancements in large Vision-Language Models have brought precise, accurate image captioning, vital for advancing multi-modal image understanding and processing. Yet these captions…
Taming OOD Actions for Offline Reinforcement Learning: An Advantage-Based Approach
Xuyang Chen, Keyu Yan, Wenhan Cao +1
Offline reinforcement learning (RL) learns policies from fixed datasets without online interactions, but suffers from distribution shift, causing inaccurate evaluation and overesti…
Frequency-Adaptive Pan-Sharpening with Mixture of Experts
Xuanhua He, Keyu Yan, Rui Li +3
Pan-sharpening involves reconstructing missing high-frequency information in multi-spectral images with low spatial resolution, using a higher-resolution panchromatic image as guid…