collaborators

7 papers

cs.LG2026

Multi-Agent Debate and Visual Information Extraction for SeePhys Pro: A 1st-Place Technical Report from ICML 2026 AI4Math Track 3 Challenge

Jiseok Kwak, Suhyeon Jo, Taewoo Kim +3

This technical report presents our approach to Challenge Track~3: SeePhys Pro at the 3rd AI for Math Workshop, where the task is to answer college-level physics questions whose sta…

cs.LG2026

Lookahead Sample Reward Guidance for Test-Time Scaling of Diffusion Models

Yeongmin Kim, Donghyeok Shin, Byeonghu Na +3

Diffusion models have demonstrated strong generative performance; however, generated samples often fail to fully align with human intent. This paper studies an efficient test-time…

cs.LG2026

Distillation of Large Language Models via Concrete Score Matching

Yeongmin Kim, Donghyeok Shin, Mina Kang +2

Large language models (LLMs) deliver remarkable performance but are costly to deploy, motivating knowledge distillation (KD) for efficient inference. Existing KD objectives typical…

cs.LG2026

AMiD: Knowledge Distillation for LLMs with -mixture Assistant Distribution

Donghyeok Shin, Yeongmin Kim, Suhyeon Jo +2

Autoregressive large language models (LLMs) have achieved remarkable improvement across many tasks but incur high computational and memory costs. Knowledge distillation (KD) mitiga…

cs.LG2026

Semantic-aware Wasserstein Policy Regularization for Large Language Model Alignment

Byeonghu Na, Hyungho Na, Yeongmin Kim +4

Large language models (LLMs) are commonly aligned with human preferences using reinforcement learning from human feedback (RLHF). In this method, LLM policies are generally optimiz…

cs.LG2025

Preference Optimization by Estimating the Ratio of the Data Distribution

Yeongmin Kim, Heesun Bae, Byeonghu Na +1

Direct preference optimization (DPO) is widely used as a simple and stable method for aligning large language models (LLMs) with human preferences. This paper investigates a genera…