activity
20242026
collaborators

9 papers

cs.CV2026

Composable Visual Tokenizers with Generator-Free Diagnostics of Learnability

Bingchen Zhao, Qiushan Guo, Ye Wang +3

We introduce CompTok, a training framework for learning visual tokenizers whose tokens are enhanced for compositionality. CompTok uses a token-conditioned diffusion decoder. By emp…

cs.CV2026

DNA: Uncovering Universal Latent Forgery Knowledge

Jingtong Dou, Chuancheng Shi, Yemin Wang +6

As generative AI achieves hyper-realism, superficial artifact detection has become obsolete. While prevailing methods rely on resource-intensive fine-tuning of black-box backbones,…

cs.CV2025

PairHuman: A High-Fidelity Photographic Dataset for Customized Dual-Person Generation

Ting Pan, Ye Wang, Peiguang Jing +3

Personalized dual-person portrait customization has considerable potential applications, such as preserving emotional memories and facilitating wedding photography planning. Howeve…

cs.CV2025

FreeControl: Efficient, Training-Free Structural Control via One-Step Attention Extraction

Jiang Lin, Xinyu Chen, Song Wu +7

Controlling the spatial and semantic structure of diffusion-generated images remains a challenge. Existing methods like ControlNet rely on handcrafted condition maps and retraining…

eess.IV2025

An Overview of the JPEG AI Learning-Based Image Coding Standard

Semih Esenlik, Yaojun Wu, Zhaobin Zhang +5

JPEG AI is an emerging learning-based image coding standard developed by Joint Photographic Experts Group (JPEG). The scope of the JPEG AI is the creation of a practical learning-b…

cs.CV2025

OmniStyle: Filtering High Quality Style Transfer Data at Scale

Ye Wang, Ruiqi Liu, Jiang Lin +4

In this paper, we introduce OmniStyle-1M, a large-scale paired style transfer dataset comprising over one million content-style-stylized image triplets across 1,000 diverse style c…