◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Aaron J. Li

University of California, Berkeley

6 papers hereh-index 4346 citations7 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author4
  • middle author2

Across the 6 of 6 papers where every author was matched, so the position is known.

fields
  • cs.CL3
  • cs.LG2
  • cs.SE1
affiliations
  • University of California, Berkeley
Homepage

identity via Semantic Scholar / OpenAlex

activity
20242026
collaborators
Showing cs.CLShow all

3 papers · 1 filter

cs.CL2026

Green Shielding: A User-Centric Approach Towards Trustworthy AI

Aaron J. Li, Nicolas Sanchez, Hao Huang +8

Large language models (LLMs) are increasingly deployed, yet their outputs can be highly sensitive to routine, non-adversarial variation in how users phrase queries, a gap not well…

cs.CL2025

Certifying LLM Safety against Adversarial Prompting

Aounon Kumar, Chirag Agarwal, Suraj Srinivas +3

Large language models (LLMs) are vulnerable to adversarial attacks that add malicious tokens to an input prompt to bypass the safety guardrails of an LLM and cause it to produce ha…

cs.CL2024

More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness

Aaron J. Li, Satyapriya Krishna, Himabindu Lakkaraju

The trustworthiness of Large Language Models (LLMs) refers to the extent to which their outputs are reliable, safe, and ethically aligned, and it has become a crucial consideration…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.