3 papers
cs.CV2023
Harnessing the Power of Multi-Lingual Datasets for Pre-training: Towards Enhancing Text Spotting Performance
Alloy Das, Sanket Biswas, Ayan Banerjee +3
The adaptation capability to a wide range of domains is crucial for scene text spotting models when deployed to real-world conditions. However, existing state-of-the-art (SOTA) app…
cs.CV2023
Diving into the Depths of Spotting Text in Multi-Domain Noisy Scenes
Alloy Das, Sanket Biswas, Umapada Pal +1
When used in a real-world noisy environment, the capacity to generalize to multiple domains is essential for any autonomous scene text spotting system. However, existing state-of-t…
cs.CV2023
GlobalDoc: A Cross-Modal Vision-Language Framework for Real-World Document Image Retrieval and Classification
Souhail Bakkali, Sanket Biswas, Zuheng Ming +4
Visual document understanding (VDU) has rapidly advanced with the development of powerful multi-modal language models. However, these models typically require extensive document pr…