Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Soft Sequence Policy Optimization
Svetlana Glazyrina, Maksim Kryzhanovskiy, Roman Ischenko
A significant portion of recent research on Large Language Model (LLM) alignment focuses on developing new policy optimization methods based on Group Relative Policy Optimization (…
cs.LG2026
Smooth Gate Functions for Soft Advantage Policy Optimization
Egor Denisov, Svetlana Glazyrina, Maksim Kryzhanovskiy +1
Group Relative Policy Optimization (GRPO) has significantly advanced the training of large language models and enhanced their reasoning capabilities, while it remains susceptible t…
cs.LG2025
Topic Modelling Black Box Optimization
Roman Akramov, Artem Khamatullin, Svetlana Glazyrina +2
Choosing the number of topics in Latent Dirichlet Allocation (LDA) is a key design decision that strongly affects both the statistical fit and interpretability of topic models.…