collaborators

5 papers

cs.LG2026

AsyncOPD: How Stale Can On-Policy Distillation Be?

Wonjun Kang, Kevin Galim, Seunghyuk Oh +9

On-policy distillation (OPD) trains a student on its own rollouts guided by teacher feedback and is becoming increasingly important for large language model (LLM) post-training. Li…

cs.CL2026

GeneralThinker: Domain-General Reasoning through Likelihood-Guided Answer-Conditioned Optimization

Shengmin Piao, Sanghyun Park

Reinforcement learning with verifiable rewards improves language model reasoning, but its reliance on domain-specific verifiers, sparse outcome rewards, and coarse-grained credit a…

cs.CL2026

SpiralThinker: Latent Reasoning through an Iterative Process with Text-Latent Interleaving

Shengmin Piao, Sanghyun Park

Recent advances in large reasoning models have been driven by reinforcement learning and test-time scaling, accompanied by growing interest in latent rather than purely textual rea…

cs.CL2026

LitE-SQL: A Lightweight and Efficient Text-to-SQL Framework with Vector-based Schema Linking and Execution-Guided Self-Correction

Shengmin Piao, Jieun Lee, Sanghyun Park

The Text-to-SQL task translates natural language questions into SQL queries, enabling intuitive database interaction for non-experts. While recent methods leveraging Large Language…

cs.CL2025

TinyThinker: Distilling Reasoning through Coarse-to-Fine Knowledge Internalization with Self-Reflection

Shengmin Piao, Sanghyun Park

Large Language Models exhibit impressive reasoning capabilities across diverse tasks, motivating efforts to distill these capabilities into smaller models through generated reasoni…