6 papers
NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap
Dasol Choi, Joonyong Park, Daegon Yu +3
We introduce NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean performance gaps arise. It comprises 15 puzzle types (25 tasks; 7,500…
Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs
Guijin Son, Seungone Kim, Catherine Arnett +73
Following the recent achievement of gold-medal performance on the IMO by frontier LLMs, the community is searching for the next meaningful and challenging target for measuring LLM…
What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models
Dasol Choi, Guijin Son, Hanwool Lee +7
Current vision-language benchmarks predominantly feature well-structured questions with clear, explicit prompts. However, real user queries are often informal and underspecified. U…
Redefining Evaluation Standards: A Unified Framework for Evaluating the Korean Capabilities of Language Models
Hanwool Lee, Dasol Choi, Sooyong Kim +6
Recent advancements in Korean large language models (LLMs) have driven numerous benchmarks and evaluation methods, yet inconsistent protocols cause up to 10 p.p performance gaps ac…
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training
Dasol Choi, Guijin Son, Soo Yong Kim +2
Visual-Language Models (VLMs) have become a powerful tool for bridging the gap between visual and linguistic understanding. However, the conventional learning approaches for VLMs o…
LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
Guijin Son, Hyunwoo Ko, Hoyoung Lee +2
LLM-as-a-Judge and reward models are widely used alternatives of multiple-choice questions or human annotators for large language model (LLM) evaluation. Their efficacy shines in e…