7 papers
Source-Grounded Data Generation for Text-to-JSON Learning
Sunghee Ahn, Guijin Son, Youngjae Yu
From financial filings to clinical records, legacy industries rely heavily on long, unstructured documents to store high-value information. Reliably extracting this information int…
K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts
Nahyun Lee, Dongkeun Yoon, Guijin Son +12
Frontier model evaluations are shifting from foundational capabilities (e.g., instruction following and reasoning) toward compositional, agentic ones, but Korean agentic benchmarks…
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…
LaRA: Layer-wise Representation Analysis for Detecting Data Contamination in RL Post-Training
Minju Gwak, Minseo Kwak, Dongseok Lee +3
Reinforcement learning (RL) post-training has shown to improve reasoning in large language models (LLMs). However, there has been little exploration on the problem of data contamin…
ResearchMath-14K: Scaling Research-Level Mathematics via Agents
Guijin Son, Seungyeop Yi, Minju Gwak +3
The frontier of mathematics is defined by problems whose solutions are not yet known, yet it remains unclear whether language models can meaningfully engage with such problems with…
Revisiting the Uniform Information Density Hypothesis in LLM Reasoning
Minju Gwak, Guijin Son, Jaehyung Kim
The Uniform Information Density (UID) hypothesis proposes that effective communication is achieved by maintaining a stable flow of information. In this work, we revisit this princi…