works on

From the 1 of 7 linked papers with an AI index.

collaborators

7 papers

cs.CL2026

Scaling Evaluation-time Compute with Reasoning Models as Evaluators

Seungone Kim, Ian Wu, Jinu Lee +8

The paper studies how using larger, chain‑of‑thought reasoning language models as evaluators—by allocating more test‑time compute—can improve the accuracy of evaluating and reranki…

cs.CL2026

K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts

Nahyun Lee, Dongkeun Yoon, Guijin Son +12

Frontier model evaluations are shifting from foundational capabilities (e.g., instruction following and reasoning) toward compositional, agentic ones, but Korean agentic benchmarks…

cs.SE2026

Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization

Anmol Agarwal, Natalie Neamtu, Pranjal Aggarwal +6

AI coding agents are increasingly used to write real-world software, but ensuring that their outputs are correct remains a fundamental challenge. Formal verification offers a promi…

cs.CL2026

On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists

Seungone Kim, Dongkeun Yoon, Kiril Gashteovski +55

With the advancement of AI capabilities, AI reviewers are beginning to be deployed in scientific peer review, yet their capability and credibility remain in question: many scientis…

cs.AI2026

Reasoning over mathematical objects: on-policy reward modeling and test time aggregation

Pranjal Aggarwal, Marjan Ghazvininejad, Seungone Kim +18

The ability to precisely derive mathematical objects is a core requirement for downstream STEM applications, including mathematics, physics, and chemistry, where reasoning must cul…

cs.CL2026

Measuring Sycophancy of Language Models in Multi-turn Dialogues

Jiseung Hong, Grace Byun, Seungone Kim +2

Large Language Models (LLMs) are expected to provide helpful and harmless responses, yet they often exhibit sycophancy--conforming to user beliefs regardless of factual accuracy or…