collaborators

5 papers

cs.LG2026

Rollout-Level Advantage-Prioritized Experience Replay for GRPO

Gyeongtae Yoo, Sanghyeok Park, Soohyuk Jang +2

Reinforcement learning from verifiable rewards with GRPO is a standard approach for post-training reasoning LLMs. It remains sample inefficient. Each rollout is used for a single g…

cs.AI2026

Metacognitive Behavioral Tuning of Large Language Models for Multi-Hop Question Answering

Ik-hwan Kim, Hyeongrok Han, Mingi Jung +5

Large Language Models (LLMs) often produce incorrect answers on multi-hop question answering even when the reasoning trace already contains a correct intermediate conclusion. We at…

cs.CL2026

Knowledge Integration Decay in Search-Augmented Reasoning of Large Language Models

Sangwon Yu, Ik-hwan Kim, Donghun Kang +6

Modern Large Language Models (LLMs) have demonstrated remarkable capabilities in complex tasks by employing search-augmented reasoning to incorporate external knowledge into long c…

cs.CL2025

Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis

Jisoo Mok, Ik-hwan Kim, Sangkwon Park +1

Personalized AI assistants, a hallmark of the human-like capabilities of Large Language Models (LLMs), are a challenging application that intertwines multiple problems in LLM resea…

cs.CL2025

Unleashing Multi-Hop Reasoning Potential in Large Language Models through Repetition of Misordered Context

Sangwon Yu, Ik-hwan Kim, Jongyoon Song +3

Multi-hop reasoning, which requires multi-step reasoning based on the supporting documents within a given context, remains challenging for large language models (LLMs). LLMs often…