1 citations · 1 across the 10 of their papers we have counts for
4 papers · 1 filter
SysAdmin: Measuring Instrumental Power-Seeking in Frontier AI
Mana Azarm, Qiyao Wei, Rahul Nambiar
Power-seeking defined as behaviors where AI systems acquire resources, evade oversight, or resist termination beyond task requirements is identified as a key driver of Loss of Cont…
Reasoning Under Pressure: How do Training Incentives Influence Chain-of-Thought Monitorability?
Matt MacDermott, Qiyao Wei, Rada Djoneva +1
AI systems that output their reasoning in natural language offer an opportunity for safety -- we can \emph{monitor} their chain of thought (CoT) for undesirable reasoning, such as…
Semantic-KG: Using Knowledge Graphs to Construct Benchmarks for Measuring Semantic Similarity
Qiyao Wei, Edward Morrell, Lea Goetz +1
Evaluating the open-form textual responses generated by Large Language Models (LLMs) typically requires measuring the semantic similarity of the response to a (human generated) ref…
Position: The ML Community Must Build an AI-Augmented Peer-Review Ecosystem
Qiyao Wei, Samuel Holt, Jing Yang +2
Peer review, the bedrock of scientific advancement in machine learning (ML), is strained by a crisis of scale. Exponential growth in manuscript submissions to premier ML venues suc…