papers

Publications (19)

cs.CV2025

Wan: Open and Advanced Large-Scale Video Generative Models

Team Wan, Ang Wang, Baole Ai +58

This report presents Wan, a comprehensive and open suite of video foundation models designed to push the boundaries of video generation. Built upon the mainstream diffusion transfo…

cs.CV2022

Panchromatic and Multispectral Image Fusion via Alternating Reverse Filtering Network

Keyu Yan, Man Zhou, Jie Huang +4

Panchromatic (PAN) and multi-spectral (MS) image fusion, named Pan-sharpening, refers to super-resolve the low-resolution (LR) multi-spectral (MS) images in the spatial domain to g…

cs.CV2026

Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoE

Haoyou Deng, Keyu Yan, Chaojie Mao +4

Mixture-of-Experts (MoE) architectures have emerged as a powerful paradigm for scaling diffusion models in visual generation. Recent advancements have focused on adaptively allocat…

cs.CV2026

DenseGRPO: From Sparse to Dense Reward for Flow Matching Model Alignment

Haoyou Deng, Keyu Yan, Chaojie Mao +4

Recent GRPO-based approaches built on flow matching models have shown remarkable improvements in human preference alignment for text-to-image generation. Nevertheless, they still s…

cs.CV2026

Perception, Verdict, and Evolution: Hindsight-Driven Self-Refining Forensics Agent for AI-Generated Image Detection

Yangjun Wu, Keyu Yan, Yu Liu +5

The rapid advancement of generative models presents a significant challenge to existing deepfake detection methods, particularly given the widespread dissemination of highly realis…

cs.CV2024

Training-Free Large Model Priors for Multiple-in-One Image Restoration

Xuanhua He, Lang Li, Yingying Wang +7

Image restoration aims to reconstruct the latent clear images from their degraded versions. Despite the notable achievement, existing methods predominantly focus on handling specif…

cs.CV2024

Pan-Mamba: Effective pan-sharpening with State Space Model

Xuanhua He, Ke Cao, Keyu Yan +4

Pan-sharpening involves integrating information from low-resolution multi-spectral and high-resolution panchromatic images to generate high-resolution multi-spectral counterparts.…

cs.CV2024

Enhancing RAW-to-sRGB with Decoupled Style Structure in Fourier Domain

Xuanhua He, Tao Hu, Guoli Wang +9

RAW to sRGB mapping, which aims to convert RAW images from smartphones into RGB form equivalent to that of Digital Single-Lens Reflex (DSLR) cameras, has become an important area o…

cs.CV2026

Wan-Image: Pushing the Boundaries of Generative Visual Intelligence

Chaojie Mao, Chen-Wei Xie, Chongyang Zhong +55

We present Wan-Image, a unified visual generation system explicitly engineered to paradigm-shift image generation models from casual synthesizers into professional-grade productivi…

cs.CV2026

Wan-Weaver: Interleaved Multi-modal Generation via Decoupled Training

Jinbo Xing, Zeyinzi Jiang, Yuxiang Tuo +15

Recent unified models have made unprecedented progress in both understanding and generation. However, while most of them accept multi-modal inputs, they typically produce only sing…

cs.CV2018

Cross-database non-frontal facial expression recognition based on transductive deep transfer learning

Keyu Yan, Wenming Zheng, Tong Zhang +2

Cross-database non-frontal expression recognition is a very meaningful but rather difficult subject in the fields of computer vision and affect computing. In this paper, we propose…

cs.CL2026

Low-resource Language Discrimination Towards Chinese Dialects with Transfer learning and Data Augmentation

Fan Xu, Yangjie Dan, Keyu Yan +2

Chinese dialects discrimination is a challenging natural language processing task due to scarce annotation resource. In this article, we develop a novel Chinese dialects discrimina…

eess.IV2022

Memory-augmented Deep Unfolding Network for Guided Image Super-resolution

Man Zhou, Keyu Yan, Jinshan Pan +3

Guided image super-resolution (GISR) aims to obtain a high-resolution (HR) target image by enhancing the spatial resolution of a low-resolution (LR) target image under the guidance…

cs.CV2024

ID-Animator: Zero-Shot Identity-Preserving Human Video Generation

Xuanhua He, Quande Liu, Shengju Qian +5

Generating high-fidelity human video with specified identities has attracted significant attention in the content generation community. However, existing techniques struggle to str…

cs.LG2026

VIPO: Value Function Inconsistency Penalized Offline Reinforcement Learning

Xuyang Chen, Keyu Yan, Guojian Wang +1

Offline reinforcement learning (RL) learns effective policies from pre-collected datasets, offering a practical solution for applications where online interactions are risky or cos…

cs.LG2026

One-Step Sampler for Boltzmann Distributions via Drifting

Wenhan Cao, Keyu Yan, Lin Zhao

We present a drifting-based framework for amortized sampling of Boltzmann distributions defined by energy functions. The method trains a one-step neural generator by projecting sam…

cs.CV2025

BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs

Zhantao Yang, Ruili Feng, Keyu Yan +13

Advancements in large Vision-Language Models have brought precise, accurate image captioning, vital for advancing multi-modal image understanding and processing. Yet these captions…

cs.LG2025

Taming OOD Actions for Offline Reinforcement Learning: An Advantage-Based Approach

Xuyang Chen, Keyu Yan, Wenhan Cao +1

Offline reinforcement learning (RL) learns policies from fixed datasets without online interactions, but suffers from distribution shift, causing inaccurate evaluation and overesti…

cs.CV2024

Frequency-Adaptive Pan-Sharpening with Mixture of Experts

Xuanhua He, Keyu Yan, Rui Li +3

Pan-sharpening involves reconstructing missing high-frequency information in multi-spectral images with low spatial resolution, using a higher-resolution panchromatic image as guid…