◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Blake Bullwinkel

15 papers hereh-index 7152 citations15 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author5
  • middle author8

Across the 13 of 15 papers where every author was matched, so the position is known.

fields
  • cs.CR6
  • cs.CL3
  • cs.LG3
  • cs.HC2
  • cs.AI1

identity via Semantic Scholar / OpenAlex

activity
20222026
most citedDEQGAN: Learning the Loss Function for PINNs with Generative Adversarial Networks

6 citations · 17 across the 12 of their papers we have counts for

collaborators
Showing cs.CLShow all

3 papers · 1 filter

cs.CL2026

Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO

Blake Bullwinkel, Eugenia Kim, Amanda Minnich +1

AI red teaming must continually adapt to evolving attackers and defenders. Reinforcement learning offers a promising approach to discovering novel attacks, and co-training methods…

cs.CL2025

Jailbreak Distillation: Renewable Safety Benchmarking

Jingyu Zhang, Ahmed Elgohary, Xiawei Wang +5

Large language models (LLMs) are rapidly deployed in critical applications, raising urgent needs for robust safety benchmarking. We propose Jailbreak Distillation (JBDistill), a no…

cs.CL2024★ 2 cited

Phi-3 Safety Post-Training: Aligning Language Models with a "Break-Fix" Cycle

Emman Haider, Daniel Perez-Becker, Thomas Portet +28

Recent innovations in language model training have demonstrated that it is possible to create highly performant models that are small enough to run on a smartphone. As these models…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.