activity
20242026
most citedWhen AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research

5 citations · 6 across the 13 of their papers we have counts for

collaborators
Showing cs.CLShow all

13 papers · 1 filter

cs.CL2026

ResearchMath-14K: Scaling Research-Level Mathematics via Agents

Guijin Son, Seungyeop Yi, Minju Gwak +3

The frontier of mathematics is defined by problems whose solutions are not yet known, yet it remains unclear whether language models can meaningfully engage with such problems with…

cs.CL2026

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs

Guijin Son, Seungone Kim, Catherine Arnett +73

Following the recent achievement of gold-medal performance on the IMO by frontier LLMs, the community is searching for the next meaningful and challenging target for measuring LLM…

cs.CL2026

KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context

Nahyun Lee, Guijin Son, Hyunwoo Ko +4

We introduce KMMMU, a native Korean benchmark for evaluating multimodal understanding in Korean cultural and institutional settings. KMMMU contains 3,466 questions from exams nativ…

cs.CL2026

Judging What We Cannot Solve: A Consequence-Based Approach for Oracle-Free Evaluation of Research-Level Math

Guijin Son, Donghun Yang, Hitesh Laxmichand Patel +5

Recent progress in reasoning models suggests that generating plausible attempts for research-level mathematics may be within reach, but verification remains a bottleneck, consuming…

cs.CL2025

Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought

Guijin Son, Donghun Yang, Hitesh Laxmichand Patel +9

Recent frontier models employ long chain-of-thought reasoning to explore solution spaces in context and achieve stonger performance. While many works study distillation to build sm…

cs.CL2025

KAIO: A Collection of More Challenging Korean Questions

Nahyun Lee, Guijin Son, Hyunwoo Ko +1

With the advancement of mid/post-training techniques, LLMs are pushing their boundaries at an accelerated pace. Legacy benchmarks saturate quickly (e.g., broad suites like MMLU ove…