◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Jordan K. Taylor

University of Queensland

3 papers hereh-index 4106 citations4 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • sole author1
  • middle author2

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.LG3
affiliations
  • University of Queensland
Homepage

identity via Semantic Scholar / OpenAlex

collaborators

3 papers

cs.LG2024

Obfuscated Activations Bypass LLM Latent-Space Defenses

Luke Bailey, Alex Serrano, Abhay Sheshadri +7

Recent latent-space monitoring techniques have shown promise as defenses against LLM attacks. These defenses act as scanners that seek to detect harmful activations before they lea…

cs.LG2024

Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning

Dan Braun, Jordan Taylor, Nicholas Goldowsky-Dill +1

Identifying the features learned by neural networks is a core challenge in mechanistic interpretability. Sparse autoencoders (SAEs), which learn a sparse, overcomplete dictionary t…

cs.LG2024

An introduction to graphical tensor notation for mechanistic interpretability

Jordan K. Taylor

Graphical tensor notation is a simple way of denoting linear operations on tensors, originating from physics. Modern deep learning consists almost entirely of operations on or betw…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.