activity
20242026
most citedOffsetBias: Leveraging Debiased Data for Tuning Evaluators

1 citations · 1 across the 4 of their papers we have counts for

collaborators

5 papers

cs.AI2026

Meta: Recursive Self-Improvement through Emergent Depth

Zae Myung Kim, Young-Jun Lee, Seungyeon Jwa +1

Self-improving LLM agents refine answers, not the process that produces those answers. Systems that add a meta-level hold that level fixed, and those that edit themselves must leav…

cs.SD2026

ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge

Jisu Jeon, Seungyeon Jwa, Joosung Lee +6

Large Audio-Language Models (LALMs) have been widely used as judge models for the automatic evaluation of generated speech. However, prior approaches predominantly focus on holisti…

cs.AI2026

RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models

San Kim, Daechul Ahn, Reokyoung Kim +3

Modern Vision-Language Models (VLMs) often struggle with strategic reasoning, i.e., anticipating and influencing other agents' actions, under uncertainty in competitive and coopera…

cs.CL2025

Becoming Experienced Judges: Selective Test-Time Learning for Evaluators

Seungyeon Jwa, Daechul Ahn, Reokyoung Kim +2

Automatic evaluation with large language models, commonly known as LLM-as-a-judge, is now standard across reasoning and alignment tasks. Despite evaluating many samples in deployme…

cs.CL20241 cited

OffsetBias: Leveraging Debiased Data for Tuning Evaluators

Junsoo Park, Seungyeon Jwa, Meiying Ren +2

Employing Large Language Models (LLMs) to assess the quality of generated responses, such as prompting instruct-tuned models or fine-tuning judge models, has become a widely adopte…