38 citations · 64 across the 23 of their papers we have counts for
3 papers · 2 filters
Leveraging Instruction Tuning and Merging for Reasoning Model Adaptation
Yu-Du Feng, Niels Mündler-Sasahara, Mark Vero +1
Reasoning language models (RLMs) have demonstrated impressive performance in domains such as mathematics and coding. These domains permit reliable verification of model outputs, wh…
Widening the Gap: Exploiting LLM Quantization via Outlier Injection
Xiaohua Zhan, Kazuki Egashira, Robin Staab +2
LLM quantization has become essential for memory-efficient deployment. Recent work has shown that quantization schemes can pose critical security risks: an adversary may release a…
Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR
Kazuki Egashira, Mark Vero, Jasper Dekoninck +3
Reinforcement Learning with Verifiable Rewards (RLVR) has become a powerful approach for improving the reasoning capabilities of large language models (LLMs). While RLVR is designe…