activity
20242026
collaborators

6 papers

cs.CV2026

InfScene-SR: Arbitrary-Size Image Super-Resolution via Iterative Joint-Denoising

Shoukun Sun, Zhe Wang, Xiang Que +2

While diffusion models have achieved state-of-the-art performance in Image Super-Resolution (SR), their prohibitive computational and memory demands restrict their training and inf…

cs.CV2026

HIME: Mitigating Object Hallucinations in LVLMs via Hallucination Insensitivity Model Editing

Ahmed Akl, Abdelwahed Khamis, Ali Cheraghian +3

Large Vision-Language Models (LVLMs) have demonstrated impressive multimodal understanding capabilities, yet they remain prone to object hallucination, where models describe non-ex…

cs.CV2025

WordCraft: Interactive Artistic Typography with Attention Awareness and Noise Blending

Zhe Wang, Jingbo Zhang, Tianyi Wei +2

Artistic typography aims to stylize input characters with visual effects that are both creative and legible. Traditional approaches rely heavily on manual design, while recent gene…

cs.LG2025

Learning-Order Autoregressive Models with Application to Molecular Graph Generation

Zhe Wang, Jiaxin Shi, Nicolas Heess +2

Autoregressive models (ARMs) have become the workhorse for sequence generation tasks, since many problems can be modeled as next-token prediction. While there appears to be a natur…

cs.LG2025

Simplified and Generalized Masked Diffusion for Discrete Data

Jiaxin Shi, Kehang Han, Zhe Wang +2

Masked (or absorbing) diffusion is actively explored as an alternative to autoregressive models for generative modeling of discrete data. However, existing work in this area has be…

cs.CV2024

Decoupled Data Augmentation for Improving Image Classification

Ruoxin Chen, Zhe Wang, Ke-Yue Zhang +5

Recent advancements in image mixing and generative data augmentation have shown promise in enhancing image classification. However, these techniques face the challenge of balancing…