most citedDocILE Benchmark for Document Information Localization and Extraction

2 citations · 2 across the 3 of their papers we have counts for

collaborators

10 papers

cs.LG2025

DocMIA: Document-Level Membership Inference Attacks against DocVQA Models

Khanh Nguyen, Raouf Kerkouche, Mario Fritz +1

Document Visual Question Answering (DocVQA) has introduced a new paradigm for end-to-end document understanding, and quickly became one of the standard benchmarks for multimodal LL…

cs.CV2024

GRIF-DM: Generation of Rich Impression Fonts using Diffusion Models

Lei Kang, Fei Yang, Kai Wang +5

Fonts are integral to creative endeavors, design processes, and artistic productions. The appropriate selection of a font can significantly enhance artwork and endow advertisements…

cs.CV2024

Comics Datasets Framework: Mix of Comics datasets for detection benchmarking

Emanuele Vivoli, Irene Campaioli, Mariateresa Nardoni +3

Comics, as a medium, uniquely combine text and images in styles often distinct from real-world visuals. For the past three decades, computational research on comics has evolved fro…

cs.CV2024

Machine Unlearning for Document Classification

Lei Kang, Mohamed Ali Souibgui, Fei Yang +3

Document understanding models have recently demonstrated remarkable performance by leveraging extensive collections of user documents. However, since documents often contain large…

cs.CV2024

Multi-Page Document Visual Question Answering using Self-Attention Scoring Mechanism

Lei Kang, Rubèn Tito, Ernest Valveny +1

Documents are 2-dimensional carriers of written communication, and as such their interpretation requires a multi-modal approach where textual and visual information are efficiently…

cs.CV2024

Multimodal Transformer for Comics Text-Cloze

Emanuele Vivoli, Joan Lafuente Baeza, Ernest Valveny Llobet +1

This work explores a closure task in comics, a medium where visual and textual elements are intricately intertwined. Specifically, Text-cloze refers to the task of selecting the co…