3 papers
cs.CL2026
CoRoVA: Compressed Representations for Vector-Augmented Code Completion
Daria Cherniuk, Nikita Sukhorukov, Danil Gusak +4
Retrieval-augmented generation has emerged as one of the most effective approaches for code completion enhancement, especially when repository-level context is important. However,…
cs.CL2026
The Chronicles of RiDiC: Generating Datasets with Controlled Popularity Distribution for Long-form Factuality Evaluation
Pavel Braslavski, Dmitrii Iarosh, Nikita Sushko +4
We present a configurable pipeline for generating multilingual sets of entities with specified characteristics, such as domain, geographical location and popularity, using data fro…
cs.CL2025
SynthDetoxM: Modern LLMs are Few-Shot Parallel Detoxification Data Annotators
Daniil Moskovskiy, Nikita Sushko, Sergey Pletenev +2
Existing approaches to multilingual text detoxification are hampered by the scarcity of parallel multilingual datasets. In this work, we introduce a pipeline for the generation of…