5 citations · 5 across the 8 of their papers we have counts for
15 papers · 1 filter
KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context
Nahyun Lee, Guijin Son, Hyunwoo Ko +4
We introduce KMMMU, a native Korean benchmark for evaluating multimodal understanding in Korean cultural and institutional settings. KMMMU contains 3,466 questions from exams nativ…
Pushing the Boundaries of Multiple Choice Evaluation to One Hundred Options
Nahyun Lee, Guijin Son
Multiple choice evaluation is widely used for benchmarking large language models, yet near ceiling accuracy in low option settings can be sustained by shortcut strategies that obsc…
Judging What We Cannot Solve: A Consequence-Based Approach for Oracle-Free Evaluation of Research-Level Math
Guijin Son, Donghun Yang, Hitesh Laxmichand Patel +5
Recent progress in reasoning models suggests that generating plausible attempts for research-level mathematics may be within reach, but verification remains a bottleneck, consuming…
Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought
Guijin Son, Donghun Yang, Hitesh Laxmichand Patel +9
Recent frontier models employ long chain-of-thought reasoning to explore solution spaces in context and achieve stonger performance. While many works study distillation to build sm…
Ko-PIQA: A Korean Physical Commonsense Reasoning Dataset with Cultural Context
Dasol Choi, Jungwhan Kim, Guijin Son
Physical commonsense reasoning datasets like PIQA are predominantly English-centric and lack cultural diversity. We introduce Ko-PIQA, a Korean physical commonsense reasoning datas…
KAIO: A Collection of More Challenging Korean Questions
Nahyun Lee, Guijin Son, Hyunwoo Ko +1
With the advancement of mid/post-training techniques, LLMs are pushing their boundaries at an accelerated pace. Legacy benchmarks saturate quickly (e.g., broad suites like MMLU ove…