activity
20242026
collaborators

5 papers

cs.CL2026

Self- and Other-Labels Induce Bidirectional Bias in LLM Judges

Songeun Chae, Min Kim, Donghoon Jung +2

As LLM-as-a-judge becomes increasingly widespread, self-preference -- the tendency of a judge to favor its own outputs -- raises growing concerns about evaluation reliability. Howe…

cs.CL2026

Anchoring LLM Gender Bias to Human Baselines: A Cross-Lingual Audit

Jiwoo Choi, Seonwoo Ahn, Tongxin Zhang +1

We audit six large language models (LLMs) for gender stereotyping across English, Korean, Chinese, and Japanese. Three were developed primarily for English-language use (Claude, GP…

cs.CL2026

Narrative Landscape: Mapping Narrative Dispositions Across LLMs

Donghoon Jung, Jiwoo Choi, Songeun Chae +1

This study proposes a quantitative framework for profiling LLM dispositions as stable, model-specific regularities in output under repeated, controlled elicitation. Using a structu…

cs.CL2025

Style over Story: Measuring LLM Narrative Preferences via Structured Selection

Donghoon Jung, Jiwoo Choi, Songeun Chae +1

We introduce a constraint-selection-based experiment design for measuring narrative preferences of Large Language Models (LLMs). This design offers an interpretable lens on LLMs' n…

cs.CL2024

A 2-step Framework for Automated Literary Translation Evaluation: Its Promises and Pitfalls

Sheikh Shafayat, Dongkeun Yoon, Woori Jang +3

In this work, we propose and evaluate the feasibility of a two-stage pipeline to evaluate literary machine translation, in a fine-grained manner, from English to Korean. The result…