Unsupervised Learning of Sentence Embeddings using Compositional n-Gram Features
arXiv:1703.02507 · doi:10.18653/v1/N18-1049
Abstract
The recent tremendous success of unsupervised word embeddings in a multitude of applications raises the obvious question if similar methods could be derived to improve embeddings (i.e. semantic representations) of word sequences as well. We present a simple but efficient unsupervised objective to train distributed representations of sentences. Our method outperforms the state-of-the-art unsupervised models on most benchmark tasks, highlighting the robustness of the produced general-purpose sentence embeddings.
NAACL 2018
References in corpus (4)
- Distributed Representations of Sentences and Documents
- Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
- A Sentimental Education: Sentiment Analysis Using Subjectivity Summarization Based on Minimum Cuts
- Charagram: Embedding Words and Sentences via Character n-grams
Cited by in corpus (123)
- Evolution of Semantic Similarity -- A Survey
- A Deep Network Model for Paraphrase Detection in Short Text Messages
- BioSentVec: creating sentence embeddings for biomedical texts
- Evaluation of sentence embeddings in downstream and linguistic probing tasks
- Learning General Purpose Distributed Sentence Representations via Large Scale Multi-task Learning
- A Review of Automated Speech and Language Features for Assessment of Cognitive and Thought Disorders
- A Survey of Malware Detection Using Deep Learning
- ESimCSE: Enhanced Sample Building Method for Contrastive Learning of Unsupervised Sentence Embedding
- Encoding word order in complex embeddings
- A Survey on Transfer Learning in Natural Language Processing
- Zero-Shot Action Recognition in Videos: A Survey
- BERT-Based Sentiment Analysis: A Software Engineering Perspective
- Multimodal Entity Linking for Tweets
- Bringing order into the realm of Transformer-based language models for artificial intelligence and law
- TNT-KID: Transformer-based Neural Tagger for Keyword Identification
- Semantic Hilbert Space for Text Representation Learning
- TSDAE: Using Transformer-based Sequential Denoising Auto-Encoder for Unsupervised Sentence Embedding Learning
- The price of debiasing automatic metrics in natural language evaluation
- PatternRank: Leveraging Pretrained Language Models and Part of Speech for Unsupervised Keyphrase Extraction
- A Multi-cascaded Model with Data Augmentation for Enhanced Paraphrase Detection in Short Texts
- SimCSE: Simple Contrastive Learning of Sentence Embeddings
- Simulated annealing for optimization of graphs and sequences
- A Tutorial on Deep Latent Variable Models of Natural Language
- Learning Semantic Textual Similarity from Conversations
- A Paragraph-level Multi-task Learning Model for Scientific Fact-Verification
- Extracting Sentence Embeddings from Pretrained Transformer Models
- Network representation learning systematic review: ancestors and current development state
- A Mathematical Exploration of Why Language Models Help Solve Downstream Tasks
- "A Passage to India": Pre-trained Word Embeddings for Indian Languages
- A Large-Scale Study of the Twitter Follower Network to Characterize the Spread of Prescription Drug Abuse Tweets
- All About Knowledge Graphs for Actions
- Filter Drug-induced Liver Injury Literature with Natural Language Processing and Ensemble Learning
- Skeleton based Zero Shot Action Recognition in Joint Pose-Language Semantic Space
- Investigating the Effects of Word Substitution Errors on Sentence Embeddings
- Tell me what you see: A zero-shot action recognition method based on natural language descriptions
- Entity Linking Meets Deep Learning: Techniques and Solutions
- Robust Cross-lingual Embeddings from Parallel Sentences
- A Blockchain-based Reliable Federated Meta-learning for Metaverse: A Dual Game Framework
- Effective Representations of Clinical Notes
- Learning Compressed Sentence Representations for On-Device Text Processing
- ReINTEL: A Multimodal Data Challenge for Responsible Information Identification on Social Network Sites
- SimCSE++: Improving Contrastive Learning for Sentence Embeddings from Two Perspectives
- Diverse Beam Search for Increased Novelty in Abstractive Summarization
- Sentence Analogies: Exploring Linguistic Relationships and Regularities in Sentence Embeddings
- Doc2Vec on the PubMed corpus: study of a new approach to generate related articles
- Correlation Coefficients and Semantic Textual Similarity
- Continual Learning for Sentence Representations Using Conceptors
- Unsupervised Learning of Sentence Representations Using Sequence Consistency
- Self-Supervised Learning for Contextualized Extractive Summarization
- A Hybrid Approach to Measure Semantic Relatedness in Biomedical Concepts
- Beyond BLEU: Training Neural Machine Translation with Semantic Similarity
- Unsupervised Paraphrasing by Simulated Annealing
- DistillCSE: Distilled Contrastive Learning for Sentence Embeddings
- Abstractive Document Summarization without Parallel Data
- Do We Need Online NLU Tools?
- Text Embeddings for Retrieval From a Large Knowledge Base
- Judge the Judges: A Large-Scale Evaluation Study of Neural Language Models for Online Review Generation
- Large-scale Hierarchical Alignment for Data-driven Text Rewriting
- Unsupervised Extraction of Phenotypes from Cancer Clinical Notes for Association Studies
- Keywords lie far from the mean of all words in local vector space
- SciLit: A Platform for Joint Scientific Literature Discovery, Summarization and Citation Generation
- SkillVet: Automated Traceability Analysis of Amazon Alexa Skills
- NLSC: Unrestricted Natural Language-based Service Composition through Sentence Embeddings
- Bayesian Paragraph Vectors
- Discourse Level Factors for Sentence Deletion in Text Simplification
- Contextual Text Embeddings for Twi
- Evaluating Multimodal Representations on Sentence Similarity: vSTS, Visual Semantic Textual Similarity Dataset
- TMIXT: A process flow for Transcribing MIXed handwritten and machine-printed Text
- Search and Learning for Unsupervised Text Generation
- Terminology-based Text Embedding for Computing Document Similarities on Technical Content
- Direct Network Transfer: Transfer Learning of Sentence Embeddings for Semantic Similarity
- Simple and Effective Paraphrastic Similarity from Parallel Translations
- GRUEN for Evaluating Linguistic Quality of Generated Text
- CBOW Is Not All You Need: Combining CBOW with the Compositional Matrix Space Model
- Structural-Aware Sentence Similarity with Recursive Optimal Transport
- Learning Malware Representation based on Execution Sequences
- GLOSS: Generative Latent Optimization of Sentence Representations
- A system for the 2019 Sentiment, Emotion and Cognitive State Task of DARPAs LORELEI project
- Question Embeddings Based on Shannon Entropy: Solving intent classification task in goal-oriented dialogue system
- Better Word Embeddings by Disentangling Contextual n-Gram Information
- Abstractive Tabular Dataset Summarization via Knowledge Base Semantic Embeddings
- NLProlog: Reasoning with Weak Unification for Question Answering in Natural Language
- Summarizing Utterances from Japanese Assembly Minutes using Political Sentence-BERT-based Method for QA Lab-PoliInfo-2 Task of NTCIR-15
- On Learning Text Style Transfer with Direct Rewards
- P-SIF: Document Embeddings Using Partition Averaging
- A reproducible experimental survey on biomedical sentence similarity: a string-based method sets the state of the art
- PAUSE: Positive and Annealed Unlabeled Sentence Embedding
- From Protocol to Screening: A Hybrid Learning Approach for Technology-Assisted Systematic Literature Reviews
- SMCLM: Semantically Meaningful Causal Language Modeling for Autoregressive Paraphrase Generation
- Detecting Logical Relation In Contract Clauses
- Bayes EMbedding (BEM): Refining Representation by Integrating Knowledge Graphs and Behavior-specific Networks
- Segmentation-free Compositional -gram Embedding
- Victim or Perpetrator? Analysis of Violent Characters Portrayals from Movie Scripts
- Bot-Match: Social Bot Detection with Recursive Nearest Neighbors Search
- Sentence transition matrix: An efficient approach that preserves sentence semantics
- Parameter-Efficient Neural Question Answering Models via Graph-Enriched Document Representations
- Abstract, Rationale, Stance: A Joint Model for Scientific Claim Verification
- Unsupervised Keyphrase Extraction by Jointly Modeling Local and Global Context
- Self-citation Analysis using Sentence Embeddings
- Vec2Sent: Probing Sentence Embeddings with Natural Language Generation
- Learning Efficient Task-Specific Meta-Embeddings with Word Prisms
- Fake News Data Collection and Classification: Iterative Query Selection for Opaque Search Engines with Pseudo Relevance Feedback
- Effective Distributed Representations for Academic Expert Search
- A Topological Approach to Compare Document Semantics Based on a New Variant of Syntactic N-grams
- Learning Interpretable and Discrete Representations with Adversarial Training for Unsupervised Text Classification
- S-APIR: News-based Business Sentiment Index
- Correlations between Word Vector Sets
- Emu: Enhancing Multilingual Sentence Embeddings with Semantic Specialization
- How Sequence-to-Sequence Models Perceive Language Styles?
- Hamming Sentence Embeddings for Information Retrieval
- Efficient Semi-Supervised Learning for Natural Language Understanding by Optimizing Diversity
- Audio Caption in a Car Setting with a Sentence-Level Loss
- Interpretable Structure-aware Document Encoders with Hierarchical Attention
- Efficient Sentence Embedding via Semantic Subspace Analysis
- STRASS: A Light and Effective Method for Extractive Summarization Based on Sentence Embeddings
- Parameter-free Sentence Embedding via Orthogonal Basis
- Testing the limits of unsupervised learning for semantic similarity
- Paraphrase Thought: Sentence Embedding Module Imitating Human Language Recognition
- Towards Understanding End-of-trip Instructions in a Taxi Ride Scenario
- Classifying Norm Conflicts using Learned Semantic Representations
- From web crawled text to project descriptions: automatic summarizing of social innovation projects
- The Effect of Downstream Classification Tasks for Evaluating Sentence Embeddings
- SAUCE: Truncated Sparse Document Signature Bit-Vectors for Fast Web-Scale Corpus Expansion