StyleRec: A Benchmark Dataset for Prompt Recovery in Writing Style Transformation
arXiv:2504.04373 · doi:10.1109/BigData62323.2024.10825143
Abstract
Prompt Recovery, reconstructing prompts from the outputs of large language models (LLMs), has grown in importance as LLMs become ubiquitous. Most users access LLMs through APIs without internal model weights, relying only on outputs and logits, which complicates recovery. This paper explores a unique prompt recovery task focused on reconstructing prompts for style transfer and rephrasing, rather than typical question-answering. We introduce a dataset created with LLM assistance, ensuring quality through multiple techniques, and test methods like zero-shot, few-shot, jailbreak, chain-of-thought, fine-tuning, and a novel canonical-prompt fallback for poor-performing cases. Our results show that one-shot and fine-tuning yield the best outcomes but highlight flaws in traditional sentence similarity metrics for evaluating prompt recovery. Contributions include (1) a benchmark dataset, (2) comprehensive experiments on prompt recovery strategies, and (3) identification of limitations in current evaluation metrics, all of which advance general prompt recovery research, where the structure of the input prompt is unrestricted.
2024 IEEE International Conference on Big Data (BigData)
References in corpus (14)
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- LoRA: Low-Rank Adaptation of Large Language Models
- Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
- Jailbroken: How Does LLM Safety Training Fail?
- The False Promise of Imitating Proprietary LLMs
- Jailbreaking Black Box Large Language Models in Twenty Queries
- Uncertainty Estimation in Autoregressive Structured Prediction
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Are You Stealing My Model? Sample Correlation for Fingerprinting Deep Neural Networks
- Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts
- Prompt Stealing Attacks Against Large Language Models
- Language Model Inversion
- Uncovering Hidden Intentions: Exploring Prompt Recovery for Deeper Insights into Generated Texts
- Was it Slander? Towards Exact Inversion of Generative Language Models