FinBERT: Financial Sentiment Analysis with Pre-trained Language Models
arXiv:1908.10063
Abstract
Financial sentiment analysis is a challenging task due to the specialized language and lack of labeled data in that domain. General-purpose models are not effective enough because of the specialized language used in a financial context. We hypothesize that pre-trained language models can help with this problem because they require fewer labeled examples and they can be further trained on domain-specific corpora. We introduce FinBERT, a language model based on BERT, to tackle NLP tasks in the financial domain. Our results show improvement in every measured metric on current state-of-the-art results for two financial sentiment analysis datasets. We find that even with a smaller training set and fine-tuning only a part of the model, FinBERT outperforms state-of-the-art machine learning methods.
This thesis is submitted in partial fulfillment for the degree of Master of Science in Information Studies: Data Science, University of Amsterdam. June 25, 2019
References in corpus (5)
- Regularizing and Optimizing LSTM Language Models
- Deep Learning: A Critical Appraisal
- Decision support from financial disclosures with deep neural networks and transfer learning
- Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books
- How to Fine-Tune BERT for Text Classification?
Cited by in corpus (25)
- GREEK-BERT: The Greeks visiting Sesame Street
- Analyzing Sustainability Reports Using Natural Language Processing
- AMMUS : A Survey of Transformer-based Pretrained Models in Natural Language Processing
- Biomedical and Clinical Language Models for Spanish: On the Benefits of Domain-Specific Pretraining in a Mid-Resource Scenario
- S&P 500 Stock Price Prediction Using Technical, Fundamental and Text Data
- Stock price prediction using BERT and GAN
- Improving Self-supervised Pre-training via a Fully-Explored Masked Language Model
- Pre-training Language Model Incorporating Domain-specific Heterogeneous Knowledge into A Unified Representation
- An Experimental Evaluation of Transformer-based Language Models in the Biomedical Domain
- Term Expansion and FinBERT fine-tuning for Hypernym and Synonym Ranking of Financial Terms
- Aspect-based Sentiment Analysis in Document -- FOMC Meeting Minutes on Economic Projection
- Boosting classification reliability of NLP transformer models in the long run
- Privacy enabled Financial Text Classification using Differential Privacy and Federated Learning
- SAUCE: Truncated Sparse Document Signature Bit-Vectors for Fast Web-Scale Corpus Expansion
- Long-term, Short-term and Sudden Event: Trading Volume Movement Prediction with Graph-based Multi-view Modeling
- Exploiting Network Structures to Improve Semantic Representation for the Financial Domain
- Customizing Contextualized Language Models forLegal Document Reviews
- On the Universality of Deep Contextual Language Models
- Detecting ESG topics using domain-specific language models and data augmentation approaches
- Interpretability in Safety-Critical FinancialTrading Systems
- Comparative Study of Language Models on Cross-Domain Data with Model Agnostic Explainability
- Question Answering over Electronic Devices: A New Benchmark Dataset and a Multi-Task Learning based QA Framework
- MDAPT: Multilingual Domain Adaptive Pretraining in a Single Model
- Tracking Turbulence Through Financial News During COVID-19
- FBERT: A Neural Transformer for Identifying Offensive Content