◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Nina Rimsky

3 papers hereh-index 51.9k citations6 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author1
  • last author1

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.CL2
  • cs.LG1

identity via Semantic Scholar / OpenAlex

collaborators

3 papers

cs.LG2024

Refusal in Language Models Is Mediated by a Single Direction

Andy Arditi, Oscar Obeso, Aaquib Syed +4

Conversational large language models are fine-tuned for both instruction-following and safety, resulting in models that obey benign requests but refuse harmful ones. While this ref…

cs.CL2024

Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

Sarah Ball, Frauke Kreuter, Nina Panickssery

Conversational large language models are trained to refuse to answer harmful questions. However, emergent jailbreaking techniques can still elicit unsafe outputs, presenting an ong…

cs.CL2024

Steering Llama 2 via Contrastive Activation Addition

Nina Panickssery, Nick Gabrieli, Julian Schulz +3

We introduce Contrastive Activation Addition (CAA), an innovative method for steering language models by modifying their activations during forward passes. CAA computes "steering v…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.