2 citations · 2 across the 3 of their papers we have counts for
10 papers
DocMIA: Document-Level Membership Inference Attacks against DocVQA Models
Khanh Nguyen, Raouf Kerkouche, Mario Fritz +1
Document Visual Question Answering (DocVQA) has introduced a new paradigm for end-to-end document understanding, and quickly became one of the standard benchmarks for multimodal LL…
GRIF-DM: Generation of Rich Impression Fonts using Diffusion Models
Lei Kang, Fei Yang, Kai Wang +5
Fonts are integral to creative endeavors, design processes, and artistic productions. The appropriate selection of a font can significantly enhance artwork and endow advertisements…
Comics Datasets Framework: Mix of Comics datasets for detection benchmarking
Emanuele Vivoli, Irene Campaioli, Mariateresa Nardoni +3
Comics, as a medium, uniquely combine text and images in styles often distinct from real-world visuals. For the past three decades, computational research on comics has evolved fro…
Machine Unlearning for Document Classification
Lei Kang, Mohamed Ali Souibgui, Fei Yang +3
Document understanding models have recently demonstrated remarkable performance by leveraging extensive collections of user documents. However, since documents often contain large…
Multi-Page Document Visual Question Answering using Self-Attention Scoring Mechanism
Lei Kang, Rubèn Tito, Ernest Valveny +1
Documents are 2-dimensional carriers of written communication, and as such their interpretation requires a multi-modal approach where textual and visual information are efficiently…
Multimodal Transformer for Comics Text-Cloze
Emanuele Vivoli, Joan Lafuente Baeza, Ernest Valveny Llobet +1
This work explores a closure task in comics, a medium where visual and textual elements are intricately intertwined. Specifically, Text-cloze refers to the task of selecting the co…