3 papers
cs.LG2026
On the Hidden Objective Biases of Group-based Reinforcement Learning
Aleksandar Fontana, Marco Simoni, Giulio Rossolini +2
Group-based reinforcement learning methods, like Group Relative Policy Optimization (GRPO), are widely used nowadays to post-train large language models. Despite their empirical su…
cs.LG2025
GTPO: Stabilizing Group Relative Policy Optimization via Gradient and Entropy Control
Marco Simoni, Aleksandar Fontana, Giulio Rossolini +2
Group Relative Policy Optimization (GRPO) is a promising policy-based approach for Large Language Model alignment, yet its performance is often limited by training instability and…
cs.AI2025
TITAN: Graph-Executable Reasoning for Cyber Threat Intelligence
Marco Simoni, Aleksandar Fontana, Andrea Saracino +1
TITAN (Threat Intelligence Through Automated Navigation) is a framework that connects natural-language cyber threat queries with executable reasoning over a structured knowledge gr…