Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024
Interleaved-Modal Chain-of-Thought
Jun Gao, Yongqi Li, Ziqiang Cao +1
Chain-of-Thought (CoT) prompting elicits large language models (LLMs) to produce a series of intermediate reasoning steps before arriving at the final answer. However, when transit…
cs.CV2024
DNTextSpotter: Arbitrary-Shaped Scene Text Spotting via Improved Denoising Training
Yu Xie, Qian Qiao, Jun Gao +5
More and more end-to-end text spotting methods based on Transformer architecture have demonstrated superior performance. These methods utilize a bipartite graph matching algorithm…