◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Avidan Shah

4 papers hereh-index 218 citations7 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2
  • middle author2

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.LG3
  • cs.CR1

identity via Semantic Scholar / OpenAlex

collaborators

4 papers

cs.LG2026

Rapid Poison: Practical Poisoning Attacks Against the Rapid Response Framework

David Huang, Jaewon Chang, Avidan Shah +2

The Rapid Response (RR) framework, deployed in production systems, including Anthropic's ASL-3 safeguards, continuously improves jailbreak-detection classifiers. When new jailbreak…

cs.CR2026

Covert Influence Between Language Models

Avidan Shah, Jay Chooi, Jinghua Ou +1

As language models increasingly consume one another's outputs, covert influence -- a phenomenon where a sender's payload (the behavioral disposition it is conditioned to propagate)…

cs.LG2026

Mitigating Adaptive Attacks against Reasoning Models with Activation Consistency Training

Avidan Shah, Jannik Brinkmann, Rico Angell

As LLMs gain stronger reasoning capabilities, their extended chain-of-thought introduces new degrees of complexity for defending against adversarial jailbreaks and prompt injection…

cs.LG2026

On-Policy Consistency Training Improves LLM Safety with Minimal Capability Degradation

Andy Han, Kristina Fujimoto, Avidan Shah +5

Aligned models can misbehave in several ways: they are often sycophantic, fall victim to jailbreaks, or fail to include appropriate safety warnings. Consistency training is a promi…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.