activity
20242026
collaborators

9 papers

cs.CV2026

LoT-Pass: Long-term-robust Image Watermarking for Image to Video Generation

Guanjie Wang, Zehua Ma, Han Fang +1

The rapid progress of image-guided video generation (I2V) has raised concerns about its potential misuse in misinformation and fraud, underscoring the urgent need for effective dig…

cs.CV2026

Flow of Truth: Proactive Temporal Forensics for Image-to-Video Generation

Yuzhuo Chen, Zehua Ma, Han Fang +3

The rapid rise of image-to-video (I2V) generation enables realistic videos to be created from a single image but also brings new forensic demands. Unlike static images, I2V content…

cs.CV2026

Allocentric Perceiver: Disentangling Allocentric Reasoning from Egocentric Visual Priors via Frame Instantiation

Hengyi Wang, Ruiqiang Zhang, Chang Liu +4

With the rising need for spatially grounded tasks such as Vision-Language Navigation/Action, allocentric perception capabilities in Vision-Language Models (VLMs) are receiving grow…

cs.CV2026

FreeText: Training-Free Text Rendering in Diffusion Transformers via Attention Localization and Spectral Glyph Injection

Ruiqiang Zhang, Hengyi Wang, Chang Liu +3

Large-scale text-to-image (T2I) diffusion models excel at open-domain synthesis but still struggle with precise text rendering, especially for multi-line layouts, dense typography,…

cs.CV2025

LAMIC: Layout-Aware Multi-Image Composition via Scalability of Multimodal Diffusion Transformer

Yuzhuo Chen, Zehua Ma, Jianhua Wang +3

In controllable image synthesis, generating coherent and consistent images from multiple references with spatial layout awareness remains an open challenge. We present LAMIC, a Lay…

cs.MM2025

TAG-WM: Tamper-Aware Generative Image Watermarking via Diffusion Inversion Sensitivity

Yuzhuo Chen, Zehua Ma, Han Fang +2

AI-generated content (AIGC) enables efficient visual creation but raises copyright and authenticity risks. As a common technique for integrity verification and source tracing, digi…