2 papers
cs.LG2025
Model Equality Testing: Which Model Is This API Serving?
Irena Gao, Percy Liang, Carlos Guestrin
Users often interact with large language models through black-box inference APIs, both for closed- and open-weight models (e.g., Llama models are popularly accessed via Amazon Bedr…
cs.CL2024
Are aligned neural networks adversarially aligned?
Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo +8
Large language models are now tuned to align with the goals of their creators, namely to be "helpful and harmless." These models should respond helpfully to user questions, but ref…