◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Amanda Askell

2 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author1

Across the 1 of 2 papers where every author was matched, so the position is known.

fields
  • cs.CL2
same name
  • Amanda Askell — 9 papers, h 18
  • Amanda Askell — 6 papers

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

most citedConstitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming

7 citations · 8 across the 2 of their papers we have counts for

collaborators
Showing cs.CLShow all

2 papers · 1 filter

cs.CL2025★ 7 cited

Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming

Mrinank Sharma, Meg Tong, Jesse Mu +40

Large language models (LLMs) are vulnerable to universal jailbreaks-prompting strategies that systematically bypass model safeguards and enable users to carry out harmful processes…

cs.CL2024★ 1 cited

Towards Implicit Bias Detection and Mitigation in Multi-Agent LLM Interactions

Angana Borah, Rada Mihalcea

As Large Language Models (LLMs) continue to evolve, they are increasingly being employed in numerous studies to simulate societies and execute diverse social tasks. However, LLMs a…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.