activity
20242026
collaborators

6 papers

cs.CV2026

Leveraging Contrastive Learning for a Similarity-Guided Tampered Document Data Generation Pipeline

Mohamed Dhouib, Davide Buscaldi, Sonia Vanier +1

Detecting tampered text in document images is a challenging task due to data scarcity. To address this, previous work has attempted to generate tampered documents using rule-based…

cs.CL2025

MUCH: A Multilingual Claim Hallucination Benchmark

Jérémie Dentan, Alexi Canesse, Davide Buscaldi +2

Claim-level Uncertainty Quantification (UQ) is a promising approach to mitigate the lack of reliability in Large Language Models (LLMs). We introduce MUCH, the first claim-level UQ…

cs.CL2025

Guess or Recall? Training CNNs to Classify and Localize Memorization in LLMs

Jérémie Dentan, Davide Buscaldi, Sonia Vanier

Verbatim memorization in Large Language Models (LLMs) is a multifaceted phenomenon involving distinct underlying mechanisms. We introduce a novel method to analyze the different fo…

cs.CL2025

Unveiling Decision-Making in LLMs for Text Classification : Extraction of influential and interpretable concepts with Sparse Autoencoders

Mathis Le Bail, Jérémie Dentan, Davide Buscaldi +1

Sparse Autoencoders (SAEs) have been successfully used to probe Large Language Models (LLMs) and extract interpretable concepts from their internal representations. These concepts…

cs.CV2025

PACT: Pruning and Clustering-Based Token Reduction for Faster Visual Language Models

Mohamed Dhouib, Davide Buscaldi, Sonia Vanier +1

Visual Language Models require substantial computational resources for inference due to the additional input tokens needed to represent visual information. However, these visual to…

cs.CR2024

Predicting memorization within Large Language Models fine-tuned for classification

Jérémie Dentan, Davide Buscaldi, Aymen Shabou +1

Large Language Models have received significant attention due to their abilities to solve a wide range of complex tasks. However these models memorize a significant proportion of t…