5 papers
UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation
Jiahao Lyu, Pei Fu, Zhenhang Li +6
In-Image Machine Translation (IIMT) aims to translate scene text in an image and render the translated text back into the original regions while preserving the overall visual appea…
IMTBench: A Multi-Scenario Cross-Modal Collaborative Evaluation Benchmark for In-Image Machine Translation
Jiahao Lyu, Pei Fu, Zhenhang Li +7
End-to-end In-Image Machine Translation (IIMT) aims to convert text embedded within an image into a target language while preserving the original visual context, layout, and render…
Class-Agnostic Region-of-Interest Matching in Document Images
Demin Zhang, Jiahao Lyu, Zhijie Shen +1
Document understanding and analysis have received a lot of attention due to their widespread application. However, existing document analysis solutions, such as document layout ana…
The Devil is in Fine-tuning and Long-tailed Problems:A New Benchmark for Scene Text Detection
Tianjiao Cao, Jiahao Lyu, Weichao Zeng +2
Scene text detection has seen the emergence of high-performing methods that excel on academic benchmarks. However, these detectors often fail to replicate such success in real-worl…
Arbitrary Reading Order Scene Text Spotter with Local Semantics Guidance
Jiahao Lyu, Wei Wang, Dongbao Yang +2
Scene text spotting has attracted the enthusiasm of relative researchers in recent years. Most existing scene text spotters follow the detection-then-recognition paradigm, where th…