2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.LG2025
It's Not You, It's Clipping: A Soft Trust-Region via Probability Smoothing for LLM RL
Madeleine Dwyer, Adam Sobey, Adriane Chapman
Training large language models (LLMs) with reinforcement learning (RL) methods such as PPO and GRPO commonly relies on ratio clipping to stabilise updates. While effective at preve…
cs.SE2024★ 2 cited
Direction for Detection: A Survey of Automated Vulnerability Detection and all of its Pain Points
Dan Ristea, Shae McFadden, Ezzeldin Shereen +4
Security vulnerabilities in software can have severe consequences; however, manual vulnerability detection is costly and does not scale, especially as agentic coding frameworks inc…