Author Identification using Multi-headed Recurrent Neural Networks
arXiv:1506.04891
Abstract
Recurrent neural networks (RNNs) are very good at modelling the flow of text, but typically need to be trained on a far larger corpus than is available for the PAN 2015 Author Identification task. This paper describes a novel approach where the output layer of a character-level RNN language model is split into several independent predictive sub-models, each representing an author, while the recurrent layer is shared by all. This allows the recurrent layer to model the language as a whole without over-fitting, while the outputs select aspects of the underlying model that reflect their author's style. The method proves competitive, ranking first in two of the four languages.
8 pages, 3 figures Version 1 was a notebook for the PAN@CLEF Author Identification challenge. Version 2 is expanded to be a full paper for CLEF2016
Cited by in corpus (7)
- Attributing Fake Images to GANs: Learning and Analyzing GAN Fingerprints
- Authorship clustering using multi-headed recurrent neural networks
- Experiments with Neural Networks for Small and Large Scale Authorship Verification
- : Author Attribute Anonymity by Adversarial Training of Neural Machine Translation
- Can You Fool AI by Doing a 180? $\unicode{x2013}$ A Case Study on Authorship Analysis of Texts by Arata Osada
- The Trumpiest Trump? Identifying a Subject's Most Characteristic Tweets
- Improving Authorship Verification using Linguistic Divergence