244 citations
- CentraleSupélecFR42 papers
- Université Paris-SaclayFR24 papers
- Département mathématiques, informatique, sciences de la donnée et technologies du numériqueFR14 papers
- Commissariat à l'Énergie Atomique et aux Énergies AlternativesFR5 papers
- University of PecsHU5 papers
- Institut de Recherche Technologique SystemXFR4 papers
- Laboratoire d'Intégration des Systèmes et des TechnologiesFR4 papers
- Airbus (France)FR3 papers
- DEDUCTEAM: Deduction modulo, interopérabilité et démonstration automatiqueFR3 papers
- Institut Gustave RoussyFR3 papers
- Institut Jean NicodFR3 papers
- Institut National de Recherche pour l'Agriculture, l'Alimentation et l'EnvironnementFR3 papers
7 papers · 2 filters
HYBRINFOX at CheckThat! 2024 -- Task 2: Enriching BERT Models with the Expert System VAGO for Subjectivity Detection
Morgane Casanova, Julien Chanson, Benjamin Icard +4
This paper presents the HYBRINFOX method used to solve Task 2 of Subjectivity detection of the CLEF 2024 CheckThat! competition. The specificity of the method is to use a hybrid sy…
A Multi-Label Dataset of French Fake News: Human and Machine Insights
Benjamin Icard, François Maine, Morgane Casanova +6
We present a corpus of 100 documents, OBSINFOX, selected from 17 sources of French press considered unreliable by expert agencies, annotated using 11 labels by 8 annotators. By col…
SaulLM-7B: A pioneering Large Language Model for Law
Pierre Colombo, Telmo Pessoa Pires, Malik Boudiaf +8
In this paper, we introduce SaulLM-7B, a large language model (LLM) tailored for the legal domain. With 7 billion parameters, SaulLM-7B is the first LLM designed explicitly for leg…
Tower: An Open Multilingual Large Language Model for Translation-Related Tasks
Duarte M. Alves, José Pombal, Nuno M. Guerreiro +10
While general-purpose large language models (LLMs) demonstrate proficiency on multiple tasks within the domain of translation, approaches based on open LLMs are competitive only wh…
Exposing propaganda: an analysis of stylistic cues comparing human annotations and machine classification
Géraud Faye, Benjamin Icard, Morgane Casanova +6
This paper investigates the language of propaganda and its stylistic features. It presents the PPN dataset, standing for Propagandist Pseudo-News, a multisource, multilingual, mult…
Towards Cross-Tokenizer Distillation: the Universal Logit Distillation Loss for LLMs
Nicolas Boizard, Kevin El Haddad, Céline Hudelot +1
Deploying large language models (LLMs) of several billion parameters can be impractical in most industrial use cases due to constraints such as cost, latency limitations, and hardw…