collaborators

7 papers

cs.CL2026

AraSSM: A bidirectional state-space encoder for Arabic masked language modeling

Ahmed Amine Aliane, Hassina Aliane, Nasredine Semmar

Pretrained Transformer encoders such as AraBERT, MARBERT, and CAMeLBERT have become the standard backbone for Arabic natural language understanding, but their self-attention mechan…

cs.CL2026

Decomposed Entailment for Factuality Checking and Hallucination Detection

Achir Oukelmoun, Nasredine Semmar, Gaël De Chalendar

The reliability of Large Language Models (LLMs) is often compromised by factual inconsistencies, including hallucinations---cases where generated content is not supported by the un…

cs.CL2026

Efficient Multilingual Neural Machine Translation via Corpus-Driven Vocabulary Pruning: An English-Arabic Case Study

Ahmed Amine Aliane, Nasredine Semmar, Hassina Aliane

The adoption of large pre-trained multilingual models for neural machine translation (MNMT) faces a major challenge: excessive memory and computational consumption due to overly la…

cs.CL2026

Injecting Structured Biomedical Knowledge into Language Models: Continual Pretraining vs. GraphRAG

Jaafer Klila, Sondes Bannour Souihi, Rahma Boujelben +2

The injection of domain-specific knowledge is crucial for adapting language models (LMs) to specialized fields such as biomedicine. While most current approaches rely on unstructur…

cs.CL2026

How DDAIR you? Disambiguated Data Augmentation for Intent Recognition

Galo Castillo-López, Alexis Lombard, Nasredine Semmar +1

Large Language Models (LLMs) are effective for data augmentation in classification tasks like intent detection. In some cases, they inadvertently produce examples that are ambiguou…

cs.LG2025

MMSE-Calibrated Few-Shot Prompting for Alzheimer's Detection

Jana Sweidan, Mounim A. El-Yacoubi, Nasredine Semmar

Prompting large language models is a training-free method for detecting Alzheimer's disease from speech transcripts. Using the ADReSS dataset, we revisit zero-shot prompting and st…