papers

Publications (24)

cs.CV2024

Diffusion-Based Hierarchical Image Steganography

Youmin Xu, Xuanyu Zhang, Jiwen Yu +3

This paper introduces Hierarchical Image Steganography, a novel method that enhances the security and capacity of embedding multiple images into a single container using diffusion…

cs.CV2024

Neural Video Fields Editing

Shuzhou Yang, Chong Mou, Jiwen Yu +3

Diffusion models have revolutionized text-driven video editing. However, applying these methods to real-world editing encounters two significant challenges: (1) the rapid increase…

cs.CV2026

MemLearner: Learning to Query Context memory for Video World Models

Jiwen Yu, Jianxiong Gao, Jianhong Bai +7

Video World Models are interactive video generation models that predict future world states based on user actions and history video frames. A critical challenge in video world mode…

cs.CV2023

CRoSS: Diffusion Model Makes Controllable, Robust and Secure Image Steganography

Jiwen Yu, Xuanyu Zhang, Youmin Xu +1

Current image steganography techniques are mainly focused on cover-based methods, which commonly have the risk of leaking secret images and poor robustness against degraded contain…

cs.CV2026

CineScene: Implicit 3D as Effective Scene Representation for Cinematic Video Generation

Kaiyi Huang, Yukun Huang, Yu Li +8

Cinematic video production requires control over scene-subject composition and camera movement, but live-action shooting remains costly due to the need for constructing physical se…

cs.CV2025

A Survey of Interactive Generative Video

Jiwen Yu, Yiran Qin, Haoxuan Che +7

Interactive Generative Video (IGV) has emerged as a crucial technology in response to the growing demand for high-quality, interactive video content across various domains. In this…

cond-mat.mtrl-sci2026

A Combined Tight Binding with Machine Learning Potential Model for Magnesium Compounds

Jiwen Yu, Arash A. Mostofi, Andrew Horsfield

We present a model for magnesium-based systems that combines density functional tight binding (DFTB) with MACE, a machine learning interatomic potential (DFTB+MACE). In this model,…

cs.CV2023

EditGuard: Versatile Image Watermarking for Tamper Localization and Copyright Protection

Xuanyu Zhang, Runyi Li, Jiwen Yu +3

In the era where AI-generated content (AIGC) models can produce stunning and lifelike images, the lingering shadow of unauthorized reproductions and malicious tampering poses immin…

cs.CV2024

V2A-Mark: Versatile Deep Visual-Audio Watermarking for Manipulation Localization and Copyright Protection

Xuanyu Zhang, Youmin Xu, Runyi Li +4

AI-generated video has revolutionized short video production, filmmaking, and personalized media, making video local editing an essential tool. However, this progress also blurs th…

cs.CV2022

GAN Prior based Null-Space Learning for Consistent Super-Resolution

Yinhuai Wang, Yujie Hu, Jiwen Yu +1

Consistency and realness have always been the two critical issues of image super-resolution. While the realness has been dramatically improved with the use of GAN prior, the state-…

cs.CV2025

Invertible Diffusion Models for Compressed Sensing

Bin Chen, Zhenyu Zhang, Weiqi Li +5

While deep neural networks (NN) significantly advance image compressed sensing (CS) by improving reconstruction quality, the necessity of training current CS NNs from scratch const…

cs.CV2023

AnimateZero: Video Diffusion Models are Zero-Shot Image Animators

Jiwen Yu, Xiaodong Cun, Chenyang Qi +4

Large-scale text-to-video (T2V) diffusion models have great progress in recent years in terms of visual quality, motion and temporal consistency. However, the generation process is…

cs.RO2026

ComSim: Building Scalable Real-World Robot Data Generation via Compositional Simulation

Yiran Qin, Jiahua Ma, Li Kang +11

Recent advancements in foundational models, such as large language models and world models, have greatly enhanced the capabilities of robotics, enabling robots to autonomously perf…

cs.CV2025

OmniX: From Unified Panoramic Generation and Perception to Graphics-Ready 3D Scenes

Yukun Huang, Jiwen Yu, Yanning Zhou +4

There are two prevalent ways for automatic 3D scene construction: procedural generation and 2D lifting. Among these, panorama-based 2D lifting has emerged as a promising technique,…

cs.CV2025

Position: Interactive Generative Video as Next-Generation Game Engine

Jiwen Yu, Yiran Qin, Haoxuan Che +5

Modern game development faces significant challenges in creativity and cost due to predetermined content in traditional game engines. Recent breakthroughs in video generation model…

cs.CV2023

DiffLLE: Diffusion-guided Domain Calibration for Unsupervised Low-light Image Enhancement

Shuzhou Yang, Xuanyu Zhang, Yinhuai Wang +3

Existing unsupervised low-light image enhancement methods lack enough effectiveness and generalization in practical applications. We suppose this is because of the absence of expli…

cs.CV2022

Zero-Shot Image Restoration Using Denoising Diffusion Null-Space Model

Yinhuai Wang, Jiwen Yu, Jian Zhang

Most existing Image Restoration (IR) models are task-specific, which can not be generalized to different degradation operators. In this work, we propose the Denoising Diffusion Nul…

cs.CV2026

MultiWorld: Scalable Multi-Agent Multi-View Video World Models

Haoyu Wu, Jiwen Yu, Yingtian Zou +1

Video world models have achieved remarkable success in simulating environmental dynamics in response to actions by users or agents. They are modeled as action-conditioned video gen…

cs.CV2025

GameFactory: Creating New Games with Generative Interactive Videos

Jiwen Yu, Yiran Qin, Xintao Wang +3

Generative videos have the potential to revolutionize game development by autonomously creating new content. In this paper, we present GameFactory, a framework for action-controlle…

cs.CV2025

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Jiwen Yu, Jianhong Bai, Yiran Qin +5

Recent advances in interactive video generation have shown promising results, yet existing approaches struggle with scene-consistent memory capabilities in long video generation du…

cs.CV2024

WorldSimBench: Towards Video Generation Models as World Simulators

Yiran Qin, Zhelun Shi, Jiwen Yu +10

Recent advancements in predictive models have demonstrated exceptional capabilities in predicting the future state of objects and scenes. However, the lack of categorization based…

cs.CV2023

FreeDoM: Training-Free Energy-Guided Conditional Diffusion Model

Jiwen Yu, Yinhuai Wang, Chen Zhao +2

Recently, conditional diffusion models have gained popularity in numerous applications due to their exceptional generation ability. However, many existing methods are training-requ…

cs.CV2025

SkillMimic: Learning Basketball Interaction Skills from Demonstrations

Yinhuai Wang, Qihan Zhao, Runyi Yu +10

Traditional reinforcement learning methods for human-object interaction (HOI) rely on labor-intensive, manually designed skill rewards that do not generalize well across different…

cs.CV2023

Unlimited-Size Diffusion Restoration

Yinhuai Wang, Jiwen Yu, Runyi Yu +1

Recently, using diffusion models for zero-shot image restoration (IR) has become a new hot paradigm. This type of method only needs to use the pre-trained off-the-shelf diffusion m…