2 papers
cs.CV2025
DNTextSpotter: Arbitrary-Shaped Scene Text Spotting via Improved Denoising Training
Yu Xie, Qian Qiao, Jun Gao +5
More and more end-to-end text spotting methods based on Transformer architecture have demonstrated superior performance. These methods utilize a bipartite graph matching algorithm…
cs.CV2024
PIP-MM: Pre-Integrating Prompt Information into Visual Encoding via Existing MLLM Structures
Tianxiang Wu, Minxin Nie, Ziqiang Cao
The Multimodal Large Language Models (MLLMs) have activated the capabilitiesof Large Language Models (LLMs) in solving visual-language tasks by integratingvisual information. The p…