Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR
Chanuk Lee, Sangwoo Park, Minki Kang +1
Reinforcement learning with verifiable rewards (RLVR) has emerged as a scalable paradigm for improving the reasoning capabilities of large language models. However, its effectivene…
cs.AI2026
PREPING: Building Agent Memory without Tasks
Yumin Choi, Sangwoo Park, Minki Kang +2
Agent memory is typically constructed either offline from curated demonstrations or online from post-deployment interactions. However, regardless of how it is built, an agent faces…
cs.AI2026
THINKSAFE: Self-Generated Safety Alignment for Reasoning Models
Seanie Lee, Sangwoo Park, Yumin Choi +6
Large reasoning models (LRMs) achieve remarkable performance by leveraging reinforcement learning (RL) on reasoning tasks to generate long chain-of-thought (CoT) reasoning. However…