1 paper · 1 filter
Md Rezwanul Haque, Md. Milon Islam, Fakhri Karray
The alignment of Small Language Models (SLMs) in the 70--500M parameter range using reinforcement learning is often considered unstable, though the underlying failure mechanisms ha…