◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

J. Rivera

3 papers hereh-index 221 citations8 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author1

Across the 2 of 3 papers where every author was matched, so the position is known.

fields
  • cs.AI2
  • cs.CL1
same name
  • J. Rivera — 2 papers, h 6

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

collaborators

3 papers

cs.AI2026

Measuring Activation Control in Large Language Models

Marek Mateusz Kowalski, Joshua Fonseca Rivera, Uzay Macar +1

Safe deployment of increasingly capable models will likely come to rely on latent-space monitoring as a complement to behavioral evaluations, especially when evaluation-aware model…

cs.AI2026

Item Response Theory for AI Safety

Joshua Fonseca Rivera, Neil Shah, David Demitri Africa +1

Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to trust and interpret, because b…

cs.CL2026

Steering Awareness: Detecting Activation Steering from Within

Joshua Fonseca Rivera, David Demitri Africa

Activation steering -- adding a vector to a model's residual stream to modify its behavior -- is widely used in safety evaluations as if the model cannot detect the intervention. W…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.