◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Sergey Berezin

3 papers hereh-index 316 citations7 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author3

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.CL1
  • cs.CR1
  • cs.LG1

identity via Semantic Scholar / OpenAlex

collaborators

3 papers

cs.LG2026

Toxicity Detection Should Measure Contextual Harm, Not Text-Intrinsic Badness

Sergei Berezin, Reza Farahbakhsh, Noel Crespi

Toxicity detection has become core safety infrastructure for online moderation, dataset filtering, and deployed language-model systems. Yet most detectors still treat toxicity as a…

cs.CL2025

Evading Toxicity Detection with ASCII-art: A Benchmark of Spatial Attacks on Moderation Systems

Sergey Berezin, Reza Farahbakhsh, Noel Crespi

We introduce a novel class of adversarial attacks on toxicity detection models that exploit language models' failure to interpret spatially structured text in the form of ASCII art…

cs.CR2025

The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs

Sergey Berezin, Reza Farahbakhsh, Noel Crespi

We present a novel class of jailbreak adversarial attacks on LLMs, termed Task-in-Prompt (TIP) attacks. Our approach embeds sequence-to-sequence tasks (e.g., cipher decoding, riddl…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.