5 papers
NAVIRA: Decoupled Stochastic Remasking for Masked Diffusion Language Models
Andrey Fomenko, Maksim Kryzhanovskiy, Svetlana Glazyrina +1
Masked diffusion language models generate text by iteratively unmasking many tokens in parallel, but this speed comes with a correction problem: tokens generated in the same step a…
Soft Sequence Policy Optimization
Svetlana Glazyrina, Maksim Kryzhanovskiy, Roman Ischenko
A significant portion of recent research on Large Language Model (LLM) alignment focuses on developing new policy optimization methods based on Group Relative Policy Optimization (…
Smooth Gate Functions for Soft Advantage Policy Optimization
Egor Denisov, Svetlana Glazyrina, Maksim Kryzhanovskiy +1
Group Relative Policy Optimization (GRPO) has significantly advanced the training of large language models and enhanced their reasoning capabilities, while it remains susceptible t…
Reinforcement Networks: novel framework for collaborative Multi-Agent Reinforcement Learning tasks
Maksim Kryzhanovskiy, Svetlana Glazyrina, Roman Ischenko +1
Modern AI systems often comprise multiple learnable components that can be naturally organized as graphs. A central challenge is the end-to-end training of such systems without res…
Topic Modelling Black Box Optimization
Roman Akramov, Artem Khamatullin, Svetlana Glazyrina +2
Choosing the number of topics in Latent Dirichlet Allocation (LDA) is a key design decision that strongly affects both the statistical fit and interpretability of topic models.…