◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Boyi Wei

Princeton University

12 papers hereh-index 9814 citations12 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author4
  • middle author8

Across the 12 of 12 papers where every author was matched, so the position is known.

fields
  • cs.CR4
  • cs.AI3
  • cs.CL3
  • cs.LG2
affiliations
  • Princeton University
Homepage
same name
  • Boyi Wei — 3 papers, h 3
  • Boyi Wei — 1 paper

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20242026
most citedSORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

6 citations · 12 across the 10 of their papers we have counts for

collaborators
Showing cs.CRShow all

4 papers · 1 filter

cs.CR2025

Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models

Boyi Wei, Zora Che, Nathaniel Li +10

Open-weight bio-foundation models present a dual-use dilemma. While holding great promise for accelerating scientific research and drug development, they could also enable bad acto…

cs.CR2025

Dynamic Risk Assessments for Offensive Cybersecurity Agents

Boyi Wei, Benedikt Stroebl, Jiacen Xu +3

Foundation models are increasingly becoming better autonomous programmers, raising the prospect that they could also automate dangerous offensive cyber-operations. Current frontier…

cs.CR2024

On Evaluating the Durability of Safeguards for Open-Weight LLMs

Xiangyu Qi, Boyi Wei, Nicholas Carlini +7

Stakeholders -- from model developers to policymakers -- seek to minimize the dual-use risks of large language models (LLMs). An open challenge to this goal is whether technical sa…

cs.CR2024

AI Risk Management Should Incorporate Both Safety and Security

Xiangyu Qi, Yangsibo Huang, Yi Zeng +22

The exposure of security vulnerabilities in safety-aligned language models, e.g., susceptibility to adversarial attacks, has shed light on the intricate interplay between AI safety…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.