6 papers
Leveraging Contrastive Learning for a Similarity-Guided Tampered Document Data Generation Pipeline
Mohamed Dhouib, Davide Buscaldi, Sonia Vanier +1
Detecting tampered text in document images is a challenging task due to data scarcity. To address this, previous work has attempted to generate tampered documents using rule-based…
MUCH: A Multilingual Claim Hallucination Benchmark
Jérémie Dentan, Alexi Canesse, Davide Buscaldi +2
Claim-level Uncertainty Quantification (UQ) is a promising approach to mitigate the lack of reliability in Large Language Models (LLMs). We introduce MUCH, the first claim-level UQ…
Guess or Recall? Training CNNs to Classify and Localize Memorization in LLMs
Jérémie Dentan, Davide Buscaldi, Sonia Vanier
Verbatim memorization in Large Language Models (LLMs) is a multifaceted phenomenon involving distinct underlying mechanisms. We introduce a novel method to analyze the different fo…
Unveiling Decision-Making in LLMs for Text Classification : Extraction of influential and interpretable concepts with Sparse Autoencoders
Mathis Le Bail, Jérémie Dentan, Davide Buscaldi +1
Sparse Autoencoders (SAEs) have been successfully used to probe Large Language Models (LLMs) and extract interpretable concepts from their internal representations. These concepts…
PACT: Pruning and Clustering-Based Token Reduction for Faster Visual Language Models
Mohamed Dhouib, Davide Buscaldi, Sonia Vanier +1
Visual Language Models require substantial computational resources for inference due to the additional input tokens needed to represent visual information. However, these visual to…
Predicting memorization within Large Language Models fine-tuned for classification
Jérémie Dentan, Davide Buscaldi, Aymen Shabou +1
Large Language Models have received significant attention due to their abilities to solve a wide range of complex tasks. However these models memorize a significant proportion of t…