Style Transfer for Texts: Retrain, Report Errors, Compare with Rewrites
arXiv:1908.06809 · doi:10.18653/v1/D19-1406
Abstract
This paper shows that standard assessment methodology for style transfer has several significant problems. First, the standard metrics for style accuracy and semantics preservation vary significantly on different re-runs. Therefore one has to report error margins for the obtained results. Second, starting with certain values of bilingual evaluation understudy (BLEU) between input and output and accuracy of the sentiment transfer the optimization of these two standard metrics diverge from the intuitive goal of the style transfer task. Finally, due to the nature of the task itself, there is a specific dependence between these two metrics that could be easily manipulated. Under these circumstances, we suggest taking BLEU between input and human-written reformulations into consideration for benchmarks. We also propose three new architectures that outperform state of the art in terms of this metric.
References in corpus (3)
Cited by in corpus (9)
- ReviewRobot: Explainable Paper Review Generation based on Knowledge Synthesis
- Style-transfer and Paraphrase: Looking for a Sensible Semantic Similarity Metric
- Improving GAN Training with Probability Ratio Clipping and Sample Reweighting
- From Theories on Styles to their Transfer in Text: Bridging the Gap with a Hierarchical Survey
- DYPLODOC: Dynamic Plots for Document Classification
- XFORMAL: A Benchmark for Multilingual Formality Style Transfer
- SE-DAE: Style-Enhanced Denoising Auto-Encoder for Unsupervised Text Style Transfer
- Evaluating the Evaluation Metrics for Style Transfer: A Case Study in Multilingual Formality Transfer
- Empirical Evaluation of Supervision Signals for Style Transfer Models