7 papers
AraSSM: A bidirectional state-space encoder for Arabic masked language modeling
Ahmed Amine Aliane, Hassina Aliane, Nasredine Semmar
Pretrained Transformer encoders such as AraBERT, MARBERT, and CAMeLBERT have become the standard backbone for Arabic natural language understanding, but their self-attention mechan…
Decomposed Entailment for Factuality Checking and Hallucination Detection
Achir Oukelmoun, Nasredine Semmar, Gaël De Chalendar
The reliability of Large Language Models (LLMs) is often compromised by factual inconsistencies, including hallucinations---cases where generated content is not supported by the un…
Efficient Multilingual Neural Machine Translation via Corpus-Driven Vocabulary Pruning: An English-Arabic Case Study
Ahmed Amine Aliane, Nasredine Semmar, Hassina Aliane
The adoption of large pre-trained multilingual models for neural machine translation (MNMT) faces a major challenge: excessive memory and computational consumption due to overly la…
Injecting Structured Biomedical Knowledge into Language Models: Continual Pretraining vs. GraphRAG
Jaafer Klila, Sondes Bannour Souihi, Rahma Boujelben +2
The injection of domain-specific knowledge is crucial for adapting language models (LMs) to specialized fields such as biomedicine. While most current approaches rely on unstructur…
How DDAIR you? Disambiguated Data Augmentation for Intent Recognition
Galo Castillo-López, Alexis Lombard, Nasredine Semmar +1
Large Language Models (LLMs) are effective for data augmentation in classification tasks like intent detection. In some cases, they inadvertently produce examples that are ambiguou…
MMSE-Calibrated Few-Shot Prompting for Alzheimer's Detection
Jana Sweidan, Mounim A. El-Yacoubi, Nasredine Semmar
Prompting large language models is a training-free method for detecting Alzheimer's disease from speech transcripts. Using the ADReSS dataset, we revisit zero-shot prompting and st…