◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Nina Panickssery

4 papers hereh-index 317 citations4 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author2
  • last author2

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.LG4

identity via Semantic Scholar / OpenAlex

collaborators

4 papers

cs.LG2026

Mechanistically Eliciting Latent Behaviors in Language Models

Andrew Mack, Nina Panickssery, Alexander Matt Turner

We aim to discover diverse, generalizable perturbations of LLM internals that can surface hidden behavioral modes. Such perturbations could help reshape model behavior and systemat…

cs.LG2025

Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs

Igor Shilov, Alex Cloud, Aryo Pradipta Gema +5

Large Language Models increasingly possess capabilities that carry dual-use risks. While data filtering has emerged as a pretraining-time mitigation, it faces significant challenge…

cs.LG2025

Mitigating Many-Shot Jailbreaking

Christopher M. Ackerman, Nina Panickssery

Many-shot jailbreaking (MSJ) is an adversarial technique that exploits the long context windows of modern LLMs to circumvent model safety training by including in the prompt many e…

cs.LG2025

Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct

Christopher Ackerman, Nina Panickssery

It has been reported that LLMs can recognize their own writing. As this has potential implications for AI safety, yet is relatively understudied, we investigate the phenomenon, see…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.