activity
20242026
collaborators

11 papers

cs.CL2026

Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation

Dongryeol Lee, Yerin Hwang, Taegwan Kang +3

While large language models (LLMs) are increasingly used as automatic judges for question answering (QA) and other reference-conditioned evaluation tasks, little is known about the…

cs.CL2026

When Wording Steers the Evaluation: Framing Bias in LLM judges

Yerin Hwang, Dongryeol Lee, Taegwan Kang +2

Large language models (LLMs) are known to produce varying responses depending on prompt phrasing, indicating that subtle guidance in phrasing can steer their answers. However, the…

cs.CL2025

Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation

Yerin Hwang, Dongryeol Lee, Kyungmin Min +3

Recently, large vision-language models (LVLMs) have emerged as the preferred tools for judging text-image alignment, yet their robustness along the visual modality remains underexp…

cs.AI2025

Program Synthesis via Test-Time Transduction

Kang-il Lee, Jahyun Koo, Seunghyun Yoon +4

We introduce transductive program synthesis, a new formulation of the program synthesis task that explicitly leverages test inputs during synthesis. While prior approaches to progr…

cs.CL2025

Can You Trick the Grader? Adversarial Persuasion of LLM Judges

Yerin Hwang, Dongryeol Lee, Taegwan Kang +2

As large language models take on growing roles as automated evaluators in practical settings, a critical question arises: Can individuals persuade an LLM judge to assign unfairly h…

cs.CL2025

Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation

Jiwon Moon, Yerin Hwang, Dongryeol Lee +3

With the growing use of large language models(LLMs) as evaluators, their application has expanded to code evaluation tasks, where they assess the correctness of generated code with…