◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

David Dobre

9 papers hereh-index 9415 citations15 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author7
  • last author1

Across the 9 of 9 papers where every author was matched, so the position is known.

fields
  • cs.LG4
  • cs.CL2
  • cs.AI1
  • gr-qc1
  • math.OC1

identity via Semantic Scholar / OpenAlex

activity
20192025
most citedEchoes from the scattering of wavepackets on wormholes

19 citations · 33 across the 6 of their papers we have counts for

collaborators
Showing cs.CLShow all

2 papers · 1 filter

cs.CL2025

A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens

David Dobre, Mehrnaz Mofakhami, Sophie Xhonneux +2

Many safety post-training methods for large language models (LLMs) are designed to modify the model's behaviour from producing unsafe answers to issuing refusals. However, such dis…

cs.CL2024

Learning diverse attacks on large language models for robust red-teaming and safety tuning

Seanie Lee, Minsu Kim, Lynn Cherif +8

Red-teaming, or identifying prompts that elicit harmful responses, is a critical step in ensuring the safe and responsible deployment of large language models (LLMs). Developing ef…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.