activity
20242026
collaborators

7 papers

cs.CV2026

SteerVTE: Seamless Video Text Editing with Style and Glyph Control

Kai Zeng, Moran Li, Zhengwei Wang +6

Visual text editing aims to precisely modify text in images and videos while preserving stylistic consistency and visual realism. Despite significant advances in the image domain,…

cs.CV2026

TextSculptor: Training and Benchmarking Scene Text Editing

Yiheng Lin, Siyu Jiao, Xiaohan Lan +12

Recent advances in Multimodal Large Language Models (MLLMs) and diffusion-based generative models have substantially improved prompt-driven image editing. However, scene text editi…

cs.CV2025

ThinkGen: Generalized Thinking for Visual Generation

Siyu Jiao, Yiheng Lin, Yujie Zhong +9

Recent progress in Multimodal Large Language Models (MLLMs) demonstrates that Chain-of-Thought (CoT) reasoning enables systematic solutions to complex understanding tasks. However,…

cs.CV2025

OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning

Shifang Zhao, Yiheng Lin, Lu Han +2

While anomaly detection has made significant progress, generating detailed analyses that incorporate industrial knowledge remains a challenge. To address this gap, we introduce Omn…

cs.CV2025

AlignGen: Boosting Personalized Image Generation with Cross-Modality Prior Alignment

Yiheng Lin, Shifang Zhao, Ting Liu +4

Personalized image generation aims to integrate user-provided concepts into text-to-image models, enabling the generation of customized content based on a given prompt. Recent zero…

cs.CV2025

DCEdit: Dual-Level Controlled Image Editing via Precisely Localized Semantics

Yihan Hu, Jianing Peng, Yiheng Lin +5

This paper presents a novel approach to improving text-guided image editing using diffusion-based models. Text-guided image editing task poses key challenge of precisly locate and…