activity
20212024
most citedTowards a Robust Detection of Language Model Generated Text: Is ChatGPT that Easy to Detect?

9 citations · 11 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CL2024

Beyond Dataset Creation: Critical View of Annotation Variation and Bias Probing of a Dataset for Online Radical Content Detection

Arij Riabi, Virginie Mouilleron, Menel Mahamdi +2

The proliferation of radical content on online platforms poses significant risks, including inciting violence and spreading extremist ideologies. Despite ongoing research, existing…

cs.CL2024

Common Ground, Diverse Roots: The Difficulty of Classifying Common Examples in Spanish Varieties

Javier A. Lopetegui, Arij Riabi, Djamé Seddah

Variations in languages across geographic regions or cultures are crucial to address to avoid biases in NLP systems designed for culturally sensitive tasks, such as hate speech det…

cs.CL2024

CamemBERT 2.0: A Smarter French Language Model Aged to Perfection

Wissam Antoun, Francis Kulumba, Rian Touchent +3

French language models, such as CamemBERT, have been widely adopted across industries for natural language processing (NLP) tasks, with models like CamemBERT seeing over 4 million…

cs.CL20239 cited

Towards a Robust Detection of Language Model Generated Text: Is ChatGPT that Easy to Detect?

Wissam Antoun, Virginie Mouilleron, Benoît Sagot +1

Recent advances in natural language processing (NLP) have led to the development of large language models (LLMs) such as ChatGPT. This paper proposes a methodology for developing a…

cs.CL2023

Data-Efficient French Language Modeling with CamemBERTa

Wissam Antoun, Benoît Sagot, Djamé Seddah

Recent advances in NLP have significantly improved the performance of language models on a variety of tasks. While these advances are largely driven by the availability of large am…

cs.CL20212 cited

Simple, Interpretable and Stable Method for Detecting Words with Usage Change across Corpora

Hila Gonen, Ganesh Jawahar, Djamé Seddah +1

The problem of comparing two bodies of text and searching for words that differ in their usage between them arises often in digital humanities and computational social science. Thi…