1 citations · 2 across the 10 of their papers we have counts for
10 papers
Model in Distress: Sentiment Analysis on French Synthetic Social Media
Pierre-Carl Langlais, Pavel Chizhov, Yannick Detrois +3
Automated analysis of customer feedback on social media is hindered by three challenges: the high cost of annotated training data, the scarcity of evaluation sets, especially in mu…
From Where Words Come: Efficient Regularization of Code Tokenizers Through Source Attribution
Pavel Chizhov, Egor Bogomolov, Ivan P. Yamshchikov
Efficiency and safety of Large Language Models (LLMs), among other factors, rely on the quality of tokenization. A good tokenizer not only improves inference speed and language und…
The Company You Keep: How LLMs Respond to Dark Triad Traits
Angelica Henestrosa, Zeyi Lu, Pavel Chizhov +1
LLMs often exhibit highly agreeable conversational styles, also known as AI sycophancy. This pattern may become problematic when interacting with user prompts that reflect negative…
Teaching Old Tokenizers New Words: Efficient Tokenizer Adaptation for Pre-trained Models
Taido Purason, Pavel Chizhov, Ivan P. Yamshchikov +1
Tokenizer adaptation plays an important role in adapting pre-trained language models to new domains or languages. In this work, we address two complementary aspects of this process…
Surface Fairness, Deep Bias: A Comparative Study of Bias in Language Models
Aleksandra Sorokovikova, Pavel Chizhov, Iuliia Eremenko +1
Modern language models are trained on large amounts of data. These data inevitably include controversial and stereotypical content, which contains all sorts of biases related to ge…
Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training
Pierre-Carl Langlais, Pavel Chizhov, Catherine Arnett +7
Large Language Models (LLMs) are pre-trained on large amounts of data from different sources and domains. Such datasets often contain trillions of tokens, including large portions…