activity
20242026
collaborators

9 papers

cs.AI2026

Position: The ML Community Must Build an AI-Augmented Peer-Review Ecosystem

Qiyao Wei, Samuel Holt, Jing Yang +2

Peer review, the bedrock of scientific advancement in machine learning (ML), is strained by a crisis of scale. Exponential growth in manuscript submissions to premier ML venues suc…

cs.AI2026

SysAdmin: Measuring Instrumental Power-Seeking in Frontier AI

Mana Azarm, Qiyao Wei, Rahul Nambiar

Power-seeking defined as behaviors where AI systems acquire resources, evade oversight, or resist termination beyond task requirements is identified as a key driver of Loss of Cont…

cs.CV2026

Paper Copilot: Tracking the Evolution of Peer Review in AI Conferences

Jing Yang, Qiyao Wei, Jiaxin Pei

The rapid growth of AI conferences is straining an already fragile peer-review system, leading to heavy reviewer workloads, expertise mismatches, inconsistent evaluation standards,…

cs.CL2025

Visualizing token importance for black-box language models

Paulius Rauba, Qiyao Wei, Mihaela van der Schaar

We consider the problem of auditing black-box large language models (LLMs) to ensure they behave reliably when deployed in production settings, particularly in high-stakes domains…

cs.AI2025

Reasoning Under Pressure: How do Training Incentives Influence Chain-of-Thought Monitorability?

Matt MacDermott, Qiyao Wei, Rada Djoneva +1

AI systems that output their reasoning in natural language offer an opportunity for safety -- we can \emph{monitor} their chain of thought (CoT) for undesirable reasoning, such as…

cs.AI2025

Semantic-KG: Using Knowledge Graphs to Construct Benchmarks for Measuring Semantic Similarity

Qiyao Wei, Edward Morrell, Lea Goetz +1

Evaluating the open-form textual responses generated by Large Language Models (LLMs) typically requires measuring the semantic similarity of the response to a (human generated) ref…