Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
The Unlearnability Phenomenon in RLVR for Language Models
Yulin Chen, He He, Chen Zhao
Reinforcement Learning with Verifiable Reward (RLVR) has proven effective in improving Large Language Model's (LLM) reasoning ability. However, the learning dynamics of RLVR remain…
cs.LG2025
Jailbreak Transferability Emerges from Shared Representations
Rico Angell, Jannik Brinkmann, He He
Jailbreak transferability is the surprising phenomenon when an adversarial attack compromising one model also elicits harmful responses from other models. Despite widespread demons…