5 papers
MemeBench: What LVLMs Miss When Interpreting Culture-Dependent Memes
Weihang Wang, Kainan Tu, Jielei Zhang +9
Large vision-language models have improved at describing visual content, but accurate descriptions do not ensure interpretation when meaning depends on knowledge beyond the pixels.…
TextFlux: An OCR-Free DiT Model for High-Fidelity Multilingual Scene Text Synthesis
Yu Xie, Jielei Zhang, Pengyu Chen +5
Diffusion-based scene text synthesis has progressed rapidly, yet existing methods commonly rely on additional visual conditioning modules and require large-scale annotated data to…
A Simple Task-aware Contrastive Local Descriptor Selection Strategy for Few-shot Learning between inter class and intra class
Qian Qiao, Yu Xie, Shaoyao Huang +1
Few-shot image classification aims to classify novel classes with few labeled samples. Recent research indicates that deep local descriptors have better representational capabiliti…
DNTextSpotter: Arbitrary-Shaped Scene Text Spotting via Improved Denoising Training
Yu Xie, Qian Qiao, Jun Gao +5
More and more end-to-end text spotting methods based on Transformer architecture have demonstrated superior performance. These methods utilize a bipartite graph matching algorithm…
TALDS-Net: Task-Aware Adaptive Local Descriptors Selection for Few-shot Image Classification
Qian Qiao, Yu Xie, Ziyin Zeng +1
Few-shot image classification aims to classify images from unseen novel classes with few samples. Recent works demonstrate that deep local descriptors exhibit enhanced representati…