activity
20242026
collaborators

10 papers

cs.CL2026

Compliance, Capability, and Conflict: Benchmarking Multimodal LLMs under System Messages

Juan Yeo, Geewook Kim

Production deployments of Multimodal Large Language Models (MLLMs) increasingly rely on system messages to govern model behavior. Yet existing benchmarks either evaluate constraint…

cs.CV2026

Do Modern Video-LLMs Need to Listen? A Benchmark Audit and Scalable Remedy

Geewook Kim, Minjoon Seo

Speech and audio encoders developed over years of community effort are routinely excluded from video understanding pipelines, not because they fail, but because benchmarks never re…

cs.CL2026

KCSAT-ML: Probing Reasoning Models with Nationwide-Cohort Human Difficulty

Sanghee Park, Geewook Kim, Kee-Eung Kim

Math reasoning benchmarks have proliferated, yet most lack a per-item difficulty signal grounded in actual human performance. We introduce KCSAT-ML, a decade (2014-2025) of Korean…

cs.CL2026

K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts

Nahyun Lee, Dongkeun Yoon, Guijin Son +12

Frontier model evaluations are shifting from foundational capabilities (e.g., instruction following and reasoning) toward compositional, agentic ones, but Korean agentic benchmarks…

cs.LG2026

Decentralized Instruction Tuning: Conflict-Aware Splitting and Weight Merging

Minsik Choi, Geewook Kim

Instruction tuning aligns large language models, including multimodal ones, with diverse user intents, but scaling to heterogeneous mixtures is hindered by gradient interference an…

cs.CV2026

State-Space Hierarchical Compression with Gated Attention and Learnable Sampling for Hour-Long Video Understanding in Large Multimodal Models

Geewook Kim, Minjoon Seo

We propose an efficient framework to compress massive video-frame features before feeding them into large multimodal models, thereby mitigating the severe token explosion arising f…