3 papers
cs.LG2026
Smooth Gate Functions for Soft Advantage Policy Optimization
Egor Denisov, Svetlana Glazyrina, Maksim Kryzhanovskiy +1
Group Relative Policy Optimization (GRPO) has significantly advanced the training of large language models and enhanced their reasoning capabilities, while it remains susceptible t…
cs.MA2025
Reinforcement Networks: novel framework for collaborative Multi-Agent Reinforcement Learning tasks
Maksim Kryzhanovskiy, Svetlana Glazyrina, Roman Ischenko +1
Modern AI systems often comprise multiple learnable components that can be naturally organized as graphs. A central challenge is the end-to-end training of such systems without res…
cs.LG2025
Topic Modelling Black Box Optimization
Roman Akramov, Artem Khamatullin, Svetlana Glazyrina +2
Choosing the number of topics in Latent Dirichlet Allocation (LDA) is a key design decision that strongly affects both the statistical fit and interpretability of topic models.…