works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.CL2026

Scaling Evaluation-time Compute with Reasoning Models as Evaluators

Seungone Kim, Ian Wu, Jinu Lee +8

The paper studies how using larger, chain‑of‑thought reasoning language models as evaluators—by allocating more test‑time compute—can improve the accuracy of evaluating and reranki…

cs.CL2026

On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists

Seungone Kim, Dongkeun Yoon, Kiril Gashteovski +55

With the advancement of AI capabilities, AI reviewers are beginning to be deployed in scientific peer review, yet their capability and credibility remain in question: many scientis…

cs.AI2026

QED-Nano: Teaching a Tiny Model to Prove Hard Theorems

LM-Provers, Yuxiao Qu, Amrith Setlur +6

Proprietary AI systems have recently demonstrated impressive capabilities on complex proof-based problems, with gold-level performance reported at the 2025 International Mathematic…

cs.AR2026

SiliconMind-V1: Multi-Agent Distillation and Debug-Reasoning Workflows for Verilog Code Generation

Mu-Chi Chen, Yu-Hung Kao, Po-Hsuan Huang +10

Large language models (LLMs) have recently emerged as a promising approach for automating Verilog code generation; however, existing methods primarily emphasize syntactic correctne…

cs.CL2025

M-Prometheus: A Suite of Open Multilingual LLM Judges

José Pombal, Dongkeun Yoon, Patrick Fernandes +5

The use of language models for automatically evaluating long-form text (LLM-as-a-judge) is becoming increasingly common, yet most LLM judges are optimized exclusively for English,…

cs.CL2025

Better Instruction-Following Through Minimum Bayes Risk

Ian Wu, Patrick Fernandes, Amanda Bertsch +3

General-purpose LLM judges capable of human-level evaluation provide not only a scalable and accurate way of evaluating instruction-following LLMs but also new avenues for supervis…