works on

From the 2 of 65 linked papers with an AI index.

activity
20242026
most citedScaling Laws for Downstream Task Performance of Large Language Models

3 citations · 5 across the 38 of their papers we have counts for

collaborators
Showing cs.AIShow all

16 papers · 1 filter

cs.AI2026

CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment

Chengxiao Wang, Enyi Jiang, Xiaojing Liao +1

Improving the safety of large language models (LLMs) often comes at the expense of utility, as globally applied safety tuning may affect model responses to both harmful and benign…

cs.AI2026

Stop Automating Peer Review Without Rigorous Evaluation

Joachim Baumann, Jiaxin Pei, Sanmi Koyejo +1

Large language models offer a tempting solution to address the peer review crisis. This position paper argues that today's AI systems should not be used to produce paper reviews. W…

cs.AI2026

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

Mubashara Akhtar, Anka Reuel, Prajna Soni +36

Artificial intelligence benchmarks are an important mechanism to measure model progress and guide deployment decisions. However, benchmarks quickly "saturate", making it difficult…

cs.AI2026

Metric Match: A Subset Selection Approach to Evaluating LLM Judge Reliability

Alyssa Unell, Natalie Dullerud, Naomi Boneh +4

LLM judges are used to reduce the need for costly human labor in evaluating open-ended text generation. However, the reliability of these judges depends critically on their alignme…

cs.AI2026

Deployment-Centered Evaluation: Predicting Query-Level Rejection Risk in a Clinical LLM System

Alyssa Unell, Miguel Fuentes, Brenna Li +4

Large language models (LLMs) are increasingly integrated into clinical systems, making it essential to evaluate the real-world utility of these systems. However, static benchmarks…

cs.AI2026

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting

Avijit Ghosh, Anka Reuel, Jenny Chim +45

AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs. The cost is interpretive: readers can…