Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning
Jean Vassoyan, Nathanaël Beau, Roman Plaud
The ability to achieve long-term goals is a key challenge in the current development of large language models (LLMs). To address this, pre-trained LLMs can be fine-tuned with reinf…
cs.CL2024
Revisiting Hierarchical Text Classification: Inference and Metrics
Roman Plaud, Matthieu Labeau, Antoine Saillenfest +1
Hierarchical text classification (HTC) is the task of assigning labels to a text within a structured space organized as a hierarchy. Recent works treat HTC as a conventional multil…