activity
20242026
most citedGoing Beyond U-Net: Assessing Vision Transformers for Semantic Segmentation in Microscopy Image Analysis

1 citations · 2 across the 10 of their papers we have counts for

collaborators

10 papers

cs.CL2026

Model in Distress: Sentiment Analysis on French Synthetic Social Media

Pierre-Carl Langlais, Pavel Chizhov, Yannick Detrois +3

Automated analysis of customer feedback on social media is hindered by three challenges: the high cost of annotated training data, the scarcity of evaluation sets, especially in mu…

cs.CL2026

From Where Words Come: Efficient Regularization of Code Tokenizers Through Source Attribution

Pavel Chizhov, Egor Bogomolov, Ivan P. Yamshchikov

Efficiency and safety of Large Language Models (LLMs), among other factors, rely on the quality of tokenization. A good tokenizer not only improves inference speed and language und…

cs.CL2026

The Company You Keep: How LLMs Respond to Dark Triad Traits

Angelica Henestrosa, Zeyi Lu, Pavel Chizhov +1

LLMs often exhibit highly agreeable conversational styles, also known as AI sycophancy. This pattern may become problematic when interacting with user prompts that reflect negative…

cs.CL2025

Teaching Old Tokenizers New Words: Efficient Tokenizer Adaptation for Pre-trained Models

Taido Purason, Pavel Chizhov, Ivan P. Yamshchikov +1

Tokenizer adaptation plays an important role in adapting pre-trained language models to new domains or languages. In this work, we address two complementary aspects of this process…

cs.CL2025

Surface Fairness, Deep Bias: A Comparative Study of Bias in Language Models

Aleksandra Sorokovikova, Pavel Chizhov, Iuliia Eremenko +1

Modern language models are trained on large amounts of data. These data inevitably include controversial and stereotypical content, which contains all sorts of biases related to ge…

cs.CL20251 cited

Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training

Pierre-Carl Langlais, Pavel Chizhov, Catherine Arnett +7

Large Language Models (LLMs) are pre-trained on large amounts of data from different sources and domains. Such datasets often contain trillions of tokens, including large portions…