activity
20172024
most citedEdit Everything: A Text-Guided Generative System for Images Editing

7 citations · 18 across the 13 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV20232 cited

Less is More: Removing Text-regions Improves CLIP Training Efficiency and Robustness

Liangliang Cao, Bowen Zhang, Chen Chen +5

The CLIP (Contrastive Language-Image Pre-training) model and its variants are becoming the de facto backbone in many applications. However, training a CLIP model from hundreds of m…

cs.CV20237 cited

Edit Everything: A Text-Guided Generative System for Images Editing

Defeng Xie, Ruichen Wang, Jian Ma +5

We introduce a new generative system called Edit Everything, which can take image and text inputs and produce image outputs. Edit Everything allows users to edit images using simpl…

cs.CV20233 cited

GlyphDraw: Seamlessly Rendering Text with Intricate Spatial Structures in Text-to-Image Generation

Jian Ma, Mingjun Zhao, Chen Chen +4

Recent breakthroughs in the field of language-guided image generation have yielded impressive achievements, enabling the creation of high-quality and diverse images based on user i…

cs.CV2023

STAIR: Learning Sparse Text and Image Representation in Grounded Tokens

Chen Chen, Bowen Zhang, Liangliang Cao +7

Image and text retrieval is one of the foundational tasks in the vision and language domain with multiple real-world applications. State-of-the-art approaches, e.g. CLIP, ALIGN, re…

cs.CV20211 cited

Visual Semantics Allow for Textual Reasoning Better in Scene Text Recognition

Yue He, Chen Chen, Jing Zhang +4

Existing Scene Text Recognition (STR) methods typically use a language model to optimize the joint probability of the 1D character sequence predicted by a visual recognition (VR) m…

cs.CV20182 cited

Rain Removal By Image Quasi-Sparsity Priors

Yinglong Wang, Shuaicheng Liu, Chen Chen +2

Rain streaks will inevitably be captured by some outdoor vision systems, which lowers the image visual quality and also interferes various computer vision applications. We present…