Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
arXiv:1908.10084
Abstract
BERT (Devlin et al., 2018) and RoBERTa (Liu et al., 2019) has set a new state-of-the-art performance on sentence-pair regression tasks like semantic textual similarity (STS). However, it requires that both sentences are fed into the network, which causes a massive computational overhead: Finding the most similar pair in a collection of 10,000 sentences requires about 50 million inference computations (~65 hours) with BERT. The construction of BERT makes it unsuitable for semantic similarity search as well as for unsupervised tasks like clustering. In this publication, we present Sentence-BERT (SBERT), a modification of the pretrained BERT network that use siamese and triplet network structures to derive semantically meaningful sentence embeddings that can be compared using cosine-similarity. This reduces the effort for finding the most similar pair from 65 hours with BERT / RoBERTa to about 5 seconds with SBERT, while maintaining the accuracy from BERT. We evaluate SBERT and SRoBERTa on common STS tasks and transfer learning tasks, where it outperforms other state-of-the-art sentence embeddings methods.
Published at EMNLP 2019
References in corpus (4)
Cited by in corpus (55)
- CLEAR: Contrastive Learning for Sentence Representation
- Entailment as Few-Shot Learner
- BertGCN: Transductive Text Classification by Combining GCN and BERT
- Self-training Improves Pre-training for Natural Language Understanding
- CO-Search: COVID-19 Information Retrieval with Semantic Search, Question Answering, and Abstractive Summarization
- On the Evaluation of Conditional GANs
- Explaining and Improving Model Behavior with k Nearest Neighbor Representations
- CLIMATE-FEVER: A Dataset for Verification of Real-World Climate Claims
- CrisisBERT: a Robust Transformer for Crisis Classification and Contextual Crisis Embedding
- Covid-Transformer: Detecting COVID-19 Trending Topics on Twitter Using Universal Sentence Encoder
- Transformer-Based Models for Question Answering on COVID19
- Self-Supervised Document Similarity Ranking via Contextualized Language Models and Hierarchical Inference
- Large-scale graph representation learning with very deep GNNs and self-supervision
- IR-BERT: Leveraging BERT for Semantic Search in Background Linking for News Articles
- Analysing the Effect of Recommendation Algorithms on the Amplification of Misinformation
- Smoothed Contrastive Learning for Unsupervised Sentence Embedding
- Supporting Clustering with Contrastive Learning
- Sentence Analogies: Exploring Linguistic Relationships and Regularities in Sentence Embeddings
- ProtoTransformer: A Meta-Learning Approach to Providing Student Feedback
- ChineseBERT: Chinese Pretraining Enhanced by Glyph and Pinyin Information
- ProtoryNet - Interpretable Text Classification Via Prototype Trajectories
- Semantic Distance: A New Metric for ASR Performance Analysis Towards Spoken Language Understanding
- MERMAID: Metaphor Generation with Symbolism and Discriminative Decoding
- Depression Status Estimation by Deep Learning based Hybrid Multi-Modal Fusion Model
- Towards Theme Detection in Personal Finance Questions
- Multitask Learning for Class-Imbalanced Discourse Classification
- Text Similarity Using Word Embeddings to Classify Misinformation
- People Still Care About Facts: Twitter Users Engage More with Factual Discourse than Misinformation--A Comparison Between COVID and General Narratives on Twitter
- ASBERT: Siamese and Triplet network embedding for open question answering
- Scaling Law for Recommendation Models: Towards General-purpose User Representations
- Paraphrase Generation as Unsupervised Machine Translation
- Multimodal Representation Learning via Maximization of Local Mutual Information
- SciSummPip: An Unsupervised Scientific Paper Summarization Pipeline
- BATS: A Spectral Biclustering Approach to Single Document Topic Modeling and Segmentation
- Reference and Document Aware Semantic Evaluation Methods for Korean Language Summarization
- Hone as You Read: A Practical Type of Interactive Summarization
- Exceeding the Limits of Visual-Linguistic Multi-Task Learning
- Detecting Escalation Level from Speech with Transfer Learning and Acoustic-Lexical Information Fusion
- A Framework for Institutional Risk Identification using Knowledge Graphs and Automated News Profiling
- Leveraging Pretrained Models for Automatic Summarization of Doctor-Patient Conversations
- Cost-effective Deployment of BERT Models in Serverless Environment
- Parameter-Efficient Neural Question Answering Models via Graph-Enriched Document Representations
- Measuring a Texts Fairness Dimensions Using Machine Learning Based on Social Psychological Factors
- Hybrid Encoder: Towards Efficient and Precise Native AdsRecommendation via Hybrid Transformer Encoding Networks
- MEDCOD: A Medically-Accurate, Emotive, Diverse, and Controllable Dialog System
- TransAug: Translate as Augmentation for Sentence Embeddings
- Latte-Mix: Measuring Sentence Semantic Similarity with Latent Categorical Mixtures
- A Framework for Learning Assessment through Multimodal Analysis of Reading Behaviour and Language Comprehension
- Can questions summarize a corpus? Using question generation for characterizing COVID-19 research
- TagRuler: Interactive Tool for Span-Level Data Programming by Demonstration
- Contrastive Semantic Similarity Learning for Image Captioning Evaluation with Intrinsic Auto-encoder
- SOCluster- Towards Intent-based Clustering of Stack Overflow Questions using Graph-Based Approach
- Self-citation Analysis using Sentence Embeddings
- On Robustness of Neural Semantic Parsers
- MIRA: Leveraging Multi-Intention Co-click Information in Web-scale Document Retrieval using Deep Neural Networks