activity
20242026
collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL2026

Compliance, Capability, and Conflict: Benchmarking Multimodal LLMs under System Messages

Juan Yeo, Geewook Kim

Production deployments of Multimodal Large Language Models (MLLMs) increasingly rely on system messages to govern model behavior. Yet existing benchmarks either evaluate constraint…

cs.CL2026

KCSAT-ML: Probing Reasoning Models with Nationwide-Cohort Human Difficulty

Sanghee Park, Geewook Kim, Kee-Eung Kim

Math reasoning benchmarks have proliferated, yet most lack a per-item difficulty signal grounded in actual human performance. We introduce KCSAT-ML, a decade (2014-2025) of Korean…

cs.CL2026

K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts

Nahyun Lee, Dongkeun Yoon, Guijin Son +12

Frontier model evaluations are shifting from foundational capabilities (e.g., instruction following and reasoning) toward compositional, agentic ones, but Korean agentic benchmarks…

cs.CL2025

MMRefine: Unveiling the Obstacles to Robust Refinement in Multimodal Large Language Models

Gio Paik, Geewook Kim, Jinbae Im

This paper introduces MMRefine, a MultiModal Refinement benchmark designed to evaluate the error refinement capabilities of Multimodal Large Language Models (MLLMs). As the emphasi…

cs.CL2025

Instruction Tuning with and without Context: Behavioral Shifts and Downstream Impact

Hyunji Lee, Seunghyun Yoon, Yunjae Won +7

Instruction tuning is a widely used approach to improve the instruction-following ability of large language models (LLMs). Instruction-tuning datasets typically include a mixture o…

cs.CL2025

Evaluating Multimodal Generative AI with Korean Educational Standards

Sanghee Park, Geewook Kim

This paper presents the Korean National Educational Test Benchmark (KoNET), a new benchmark designed to evaluate Multimodal Generative AI Systems using Korean national educational…