9 citations · 11 across the 6 of their papers we have counts for
6 papers
Beyond Dataset Creation: Critical View of Annotation Variation and Bias Probing of a Dataset for Online Radical Content Detection
Arij Riabi, Virginie Mouilleron, Menel Mahamdi +2
The proliferation of radical content on online platforms poses significant risks, including inciting violence and spreading extremist ideologies. Despite ongoing research, existing…
Common Ground, Diverse Roots: The Difficulty of Classifying Common Examples in Spanish Varieties
Javier A. Lopetegui, Arij Riabi, Djamé Seddah
Variations in languages across geographic regions or cultures are crucial to address to avoid biases in NLP systems designed for culturally sensitive tasks, such as hate speech det…
CamemBERT 2.0: A Smarter French Language Model Aged to Perfection
Wissam Antoun, Francis Kulumba, Rian Touchent +3
French language models, such as CamemBERT, have been widely adopted across industries for natural language processing (NLP) tasks, with models like CamemBERT seeing over 4 million…
Towards a Robust Detection of Language Model Generated Text: Is ChatGPT that Easy to Detect?
Wissam Antoun, Virginie Mouilleron, Benoît Sagot +1
Recent advances in natural language processing (NLP) have led to the development of large language models (LLMs) such as ChatGPT. This paper proposes a methodology for developing a…
Data-Efficient French Language Modeling with CamemBERTa
Wissam Antoun, Benoît Sagot, Djamé Seddah
Recent advances in NLP have significantly improved the performance of language models on a variety of tasks. While these advances are largely driven by the availability of large am…
Simple, Interpretable and Stable Method for Detecting Words with Usage Change across Corpora
Hila Gonen, Ganesh Jawahar, Djamé Seddah +1
The problem of comparing two bodies of text and searching for words that differ in their usage between them arises often in digital humanities and computational social science. Thi…