collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

FireRed-OCR Technical Report

Hao Wu, Haoran Lou, Xinyue Li +19

We present FireRed-OCR, a systematic framework to specialize general VLMs into high-performance OCR models. Large Vision-Language Models (VLMs) have demonstrated impressive general…

cs.CV2026

IdGlow: Dynamic Identity Modulation for Multi-Subject Generation

Honghao Cai, Xiangyuan Wang, Jing Li +15

Multi-subject image generation requires seamlessly harmonizing multiple reference identities within a coherent scene. However, existing methods relying on rigid spatial masks or lo…

cs.CV2024

Target-Driven Distillation: Consistency Distillation with Target Timestep Selection and Decoupled Guidance

Cunzheng Wang, Ziyuan Guo, Yuxuan Duan +4

Consistency distillation methods have demonstrated significant success in accelerating generative tasks of diffusion models. However, since previous consistency distillation method…

cs.CV2024

StableGarment: Garment-Centric Generation via Stable Diffusion

Rui Wang, Hailong Guo, Jiaming Liu +6

In this paper, we introduce StableGarment, a unified framework to tackle garment-centric(GC) generation tasks, including GC text-to-image, controllable GC text-to-image, stylized G…

cs.CV2024

Stable-Makeup: When Real-World Makeup Transfer Meets Diffusion Model

Yuxuan Zhang, Yirui Yuan, Yiren Song +1

Current makeup transfer methods are limited to simple makeup styles, making them difficult to apply in real-world scenarios. In this paper, we introduce Stable-Makeup, a novel diff…

cs.CV202424 cited

InstantID: Zero-shot Identity-Preserving Generation in Seconds

Qixun Wang, Xu Bai, Haofan Wang +5

There has been significant progress in personalized image synthesis with methods such as Textual Inversion, DreamBooth, and LoRA. Yet, their real-world applicability is hindered by…