554 citations · 1k across the 21 of their papers we have counts for
4 papers · 1 filter
FashionLOGO: Prompting Multimodal Large Language Models for Fashion Logo Embeddings
Zhen Wang, Da Li, Yulin Su +3
Logo embedding models convert the product logos in images into vectors, enabling their utilization for logo recognition and detection within e-commerce platforms. This facilitates…
Kaleido-BERT: Vision-Language Pre-training on Fashion Domain
Mingchen Zhuge, Dehong Gao, Deng-Ping Fan +5
We present a new vision-language (VL) pre-training model dubbed Kaleido-BERT, which introduces a novel kaleido strategy for fashion cross-modality representations from transformers…
One-shot Text Field Labeling using Attention and Belief Propagation for Structure Information Extraction
Mengli Cheng, Minghui Qiu, Xing Shi +2
Structured information extraction from document images usually consists of three steps: text detection, text recognition, and text field labeling. While text detection and text rec…
IncepText: A New Inception-Text Module with Deformable PSROI Pooling for Multi-Oriented Scene Text Detection
Qiangpeng Yang, Mengli Cheng, Wenmeng Zhou +4
Incidental scene text detection, especially for multi-oriented text regions, is one of the most challenging tasks in many computer vision applications. Different from the common ob…