2 papers
cs.LG2026
Value-Gradient Hypothesis of RL for LLMs
Arip Asadulaev, Daniil Ognev, Karim Salta +1
Reinforcement learning substantially improves pretrained language models, but it remains understudied why critic-free methods such as PPO and GRPO work as well as they do, and when…
cs.LG2025
A Granular Grassmannian Clustering Framework via the Schubert Variety of Best Fit
Karim Salta, Michael Kirby, Chris Peterson
In many classification and clustering tasks, it is useful to compute a geometric representative for a dataset or a cluster, such as a mean or median. When datasets are represented…