activity
20242026
most citedFreqMamba: Viewing Mamba from a Frequency Perspective for Image Deraining

1 citations · 2 across the 13 of their papers we have counts for

collaborators
Showing cs.CVShow all

12 papers · 1 filter

cs.CV2026

Style as Cover: Deep Image Steganography via Stylized Transmission

Qi Li, Jidong Yang, Huaike Yu +6

Image steganography hides secret message within normal images, with most existing works relying on cover-preserving transmission. However, such a paradigm becomes vulnerable once t…

cs.CV2026

One Prompt Is Enough: Watermark Laundering Through Foundation Image Models

Jidong Yang, Qi Li, Wei Zong +5

Invisible watermarks are typically evaluated against predefined perturbations such as compression, blur, noise, cropping, and denoising. Public foundation image models expose a dis…

cs.CV2026

AnchorEdit: Maintaining Temporal Consistency in Multi-turn Image Editing via Causal Memory

Hang Xu, Xiaoxiao Ma, Guohui Zhang +7

Multi-turn image editing is essential for iterative design, yet current models often struggle with identity drift and error accumulation over successive steps. While existing resea…

cs.CV2026

OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation

Guohui Zhang, XiaoXiao Ma, Jie Huang +9

Recent advances in joint audio-video generation have been remarkable, yet real-world applications demand strong per-modality fidelity, cross-modal alignment, and fine-grained synch…

cs.CV2025

MaskFocus: Focusing Policy Optimization on Critical Steps for Masked Image Generation

Guohui Zhang, Hu Yu, Xiaoxiao Ma +3

Reinforcement learning (RL) has demonstrated significant potential for post-training language models and autoregressive visual generative models, but adapting RL to masked generati…

cs.CV2025

Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy

Xiaoxiao Ma, Feng Zhao, Pengyang Ling +6

In this work, we first revisit the sampling issues in current autoregressive (AR) image generation models and identify that image tokens, unlike text tokens, exhibit lower informat…