17 papers
Language Shapes Instruction Hierarchy Compliance in Multilingual LLMs
Jiwon Moon, Yerin Hwang, Kyomin Jung
Instruction hierarchy (IH) requires models to prioritize instructions by source, ensuring that higher-priority instructions override lower-priority ones. Despite its importance for…
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
Dongryeol Lee, Yerin Hwang, Taegwan Kang +3
While large language models (LLMs) are increasingly used as automatic judges for question answering (QA) and other reference-conditioned evaluation tasks, little is known about the…
Casual as an Anchor: Resolving Supervision Misalignment in Formality Transfer Dataset
Hyojeong Yu, Hyukhun Koh, Minsung Kim +1
Formality transfer is commonly framed as a symmetric bidirectional task between informal and formal registers. We argue that this framing conceals a supervision design flaw in exis…
When Wording Steers the Evaluation: Framing Bias in LLM judges
Yerin Hwang, Dongryeol Lee, Taegwan Kang +2
Large language models (LLMs) are known to produce varying responses depending on prompt phrasing, indicating that subtle guidance in phrasing can steer their answers. However, the…
Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation
Yerin Hwang, Dongryeol Lee, Kyungmin Min +3
Recently, large vision-language models (LVLMs) have emerged as the preferred tools for judging text-image alignment, yet their robustness along the visual modality remains underexp…
Program Synthesis via Test-Time Transduction
Kang-il Lee, Jahyun Koo, Seunghyun Yoon +4
We introduce transductive program synthesis, a new formulation of the program synthesis task that explicitly leverages test inputs during synthesis. While prior approaches to progr…