activity
20202026
collaborators

5 papers

cs.CV2026

Leveraging Contrastive Learning for a Similarity-Guided Tampered Document Data Generation Pipeline

Mohamed Dhouib, Davide Buscaldi, Sonia Vanier +1

Detecting tampered text in document images is a challenging task due to data scarcity. To address this, previous work has attempted to generate tampered documents using rule-based…

cs.CL2025

MUCH: A Multilingual Claim Hallucination Benchmark

Jérémie Dentan, Alexi Canesse, Davide Buscaldi +2

Claim-level Uncertainty Quantification (UQ) is a promising approach to mitigate the lack of reliability in Large Language Models (LLMs). We introduce MUCH, the first claim-level UQ…

cs.CV2025

PACT: Pruning and Clustering-Based Token Reduction for Faster Visual Language Models

Mohamed Dhouib, Davide Buscaldi, Sonia Vanier +1

Visual Language Models require substantial computational resources for inference due to the additional input tokens needed to represent visual information. However, these visual to…

cs.CR2024

Predicting memorization within Large Language Models fine-tuned for classification

Jérémie Dentan, Davide Buscaldi, Aymen Shabou +1

Large Language Models have received significant attention due to their abilities to solve a wide range of complex tasks. However these models memorize a significant proportion of t…

cs.CV2020

VisualWordGrid: Information Extraction From Scanned Documents Using A Multimodal Approach

Mohamed Kerroumi, Othmane Sayem, Aymen Shabou

We introduce a novel approach for scanned document representation to perform field extraction. It allows the simultaneous encoding of the textual, visual and layout information in…