collaborators

6 papers

cs.CL2026

NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap

Dasol Choi, Joonyong Park, Daegon Yu +3

We introduce NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean performance gaps arise. It comprises 15 puzzle types (25 tasks; 7,500…

cs.CL2026

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs

Guijin Son, Seungone Kim, Catherine Arnett +73

Following the recent achievement of gold-medal performance on the IMO by frontier LLMs, the community is searching for the next meaningful and challenging target for measuring LLM…

cs.CV2026

What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models

Dasol Choi, Guijin Son, Hanwool Lee +7

Current vision-language benchmarks predominantly feature well-structured questions with clear, explicit prompts. However, real user queries are often informal and underspecified. U…

cs.CE2026

Redefining Evaluation Standards: A Unified Framework for Evaluating the Korean Capabilities of Language Models

Hanwool Lee, Dasol Choi, Sooyong Kim +6

Recent advancements in Korean large language models (LLMs) have driven numerous benchmarks and evaluation methods, yet inconsistent protocols cause up to 10 p.p performance gaps ac…

cs.CL2024

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training

Dasol Choi, Guijin Son, Soo Yong Kim +2

Visual-Language Models (VLMs) have become a powerful tool for bridging the gap between visual and linguistic understanding. However, the conventional learning approaches for VLMs o…

cs.CL2024

LLM-as-a-Judge & Reward Model: What They Can and Cannot Do

Guijin Son, Hyunwoo Ko, Hoyoung Lee +2

LLM-as-a-Judge and reward models are widely used alternatives of multiple-choice questions or human annotators for large language model (LLM) evaluation. Their efficacy shines in e…