6 papers · 1 filter
Reading in the Dark: Low-light Scene Text Recognition
Xuanshuo Fu, Lei Kang, Ernest Valveny +2
Accurate text recognition in low-light environments is essential for intelligent systems in applications ranging from autonomous vehicles to smart surveillance. However, challenges…
Learning to Decipher from Pixels: A Case Study of Copiale
Lei Kang, Giuseppe De Gregorio, Raphaela Heil +2
Historical encrypted manuscripts require both paleographic interpretation of cipher symbols and cryptanalytic recovery of plaintext. Most existing computational workflows rely on a…
AVIR: Adaptive Visual In-Document Retrieval for Efficient Multi-Page Document Question Answering
Zongmin Li, Yachuan Li, Lei Kang +2
Multi-page Document Visual Question Answering (MP-DocVQA) remains challenging because long documents not only strain computational resources but also reduce the effectiveness of th…
Machine Unlearning for Document Classification
Lei Kang, Mohamed Ali Souibgui, Fei Yang +3
Document understanding models have recently demonstrated remarkable performance by leveraging extensive collections of user documents. However, since documents often contain large…
Multi-Page Document Visual Question Answering using Self-Attention Scoring Mechanism
Lei Kang, Rubèn Tito, Ernest Valveny +1
Documents are 2-dimensional carriers of written communication, and as such their interpretation requires a multi-modal approach where textual and visual information are efficiently…
Privacy-Aware Document Visual Question Answering
Rubèn Tito, Khanh Nguyen, Marlon Tobaben +11
Document Visual Question Answering (DocVQA) has quickly grown into a central task of document understanding. But despite the fact that documents contain sensitive or copyrighted in…