1 paper · 1 filter
Nicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham +11
We introduce the first model-stealing attack that extracts precise, nontrivial information from black-box production language models like OpenAI's ChatGPT or Google's PaLM-2. Speci…