activity
20242026
most citedWhen AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research

5 citations · 5 across the 8 of their papers we have counts for

collaborators
Showing cs.CLShow all

15 papers · 1 filter

cs.CL2026

KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context

Nahyun Lee, Guijin Son, Hyunwoo Ko +4

We introduce KMMMU, a native Korean benchmark for evaluating multimodal understanding in Korean cultural and institutional settings. KMMMU contains 3,466 questions from exams nativ…

cs.CL2026

Pushing the Boundaries of Multiple Choice Evaluation to One Hundred Options

Nahyun Lee, Guijin Son

Multiple choice evaluation is widely used for benchmarking large language models, yet near ceiling accuracy in low option settings can be sustained by shortcut strategies that obsc…

cs.CL2026

Judging What We Cannot Solve: A Consequence-Based Approach for Oracle-Free Evaluation of Research-Level Math

Guijin Son, Donghun Yang, Hitesh Laxmichand Patel +5

Recent progress in reasoning models suggests that generating plausible attempts for research-level mathematics may be within reach, but verification remains a bottleneck, consuming…

cs.CL2025

Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought

Guijin Son, Donghun Yang, Hitesh Laxmichand Patel +9

Recent frontier models employ long chain-of-thought reasoning to explore solution spaces in context and achieve stonger performance. While many works study distillation to build sm…

cs.CL2025

Ko-PIQA: A Korean Physical Commonsense Reasoning Dataset with Cultural Context

Dasol Choi, Jungwhan Kim, Guijin Son

Physical commonsense reasoning datasets like PIQA are predominantly English-centric and lack cultural diversity. We introduce Ko-PIQA, a Korean physical commonsense reasoning datas…

cs.CL2025

KAIO: A Collection of More Challenging Korean Questions

Nahyun Lee, Guijin Son, Hyunwoo Ko +1

With the advancement of mid/post-training techniques, LLMs are pushing their boundaries at an accelerated pace. Legacy benchmarks saturate quickly (e.g., broad suites like MMLU ove…