XLNet: Generalized Autoregressive Pretraining for Language Understanding
arXiv:1906.08237
Abstract
With the capability of modeling bidirectional contexts, denoising autoencoding based pretraining like BERT achieves better performance than pretraining approaches based on autoregressive language modeling. However, relying on corrupting the input with masks, BERT neglects dependency between the masked positions and suffers from a pretrain-finetune discrepancy. In light of these pros and cons, we propose XLNet, a generalized autoregressive pretraining method that (1) enables learning bidirectional contexts by maximizing the expected likelihood over all permutations of the factorization order and (2) overcomes the limitations of BERT thanks to its autoregressive formulation. Furthermore, XLNet integrates ideas from Transformer-XL, the state-of-the-art autoregressive model, into pretraining. Empirically, under comparable experiment settings, XLNet outperforms BERT on 20 tasks, often by a large margin, including question answering, natural language inference, sentiment analysis, and document ranking.
Pretrained models and code are available at https://github.com/zihangdai/xlnet
References in corpus (13)
- Unsupervised Data Augmentation for Consistency Training
- ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
- A Deep Relevance Matching Model for Ad-hoc Retrieval
- Snorkel: Rapid Training Data Creation with Weak Supervision
- End-to-End Neural Ad-hoc Ranking with Kernel Pooling
- MADE: Masked Autoencoder for Distribution Estimation
- Multi-Task Deep Neural Networks for Natural Language Understanding
- Word-Entity Duet Representations for Document Ranking
- A Surprisingly Robust Trick for Winograd Schema Challenge
- NewsQA: A Machine Comprehension Dataset
- Option Comparison Network for Multiple-choice Reading Comprehension
- Dual Co-Matching Network for Multi-choice Reading Comprehension
- Improving Question Answering with External Knowledge
Cited by in corpus (344)
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- Attention Mechanisms in Computer Vision: A Survey
- BERTScore: Evaluating Text Generation with BERT
- WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
- Pre-trained Models for Natural Language Processing: A Survey
- ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- ChatGPT and a New Academic Reality: Artificial Intelligence-Written Research Papers and the Ethics of the Large Language Models in Scholarly Publishing
- VL-BERT: Pre-training of Generic Visual-Linguistic Representations
- Multilingual Denoising Pre-training for Neural Machine Translation
- Q8BERT: Quantized 8Bit BERT
- Well-Read Students Learn Better: On the Importance of Pre-training Compact Models
- TERA: Self-Supervised Learning of Transformer Encoder Representation for Speech
- Reducing Transformer Depth on Demand with Structured Dropout
- VLP: A Survey on Vision-Language Pre-training
- CodeBERT: A Pre-Trained Model for Programming and Natural Languages
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- DocBERT: BERT for Document Classification
- Controllable Protein Design with Language Models
- ERNIE 3.0: Large-scale Knowledge Enhanced Pre-training for Language Understanding and Generation
- Portuguese Named Entity Recognition using BERT-CRF
- Incorporating BERT into Neural Machine Translation
- SMILES Transformer: Pre-trained Molecular Fingerprint for Low Data Drug Discovery
- Recent Advances in Natural Language Processing via Large Pre-Trained Language Models: A Survey
- GraphCodeBERT: Pre-training Code Representations with Data Flow
- A Survey on Long-Tailed Visual Recognition
- ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data
- Retrieving and Reading: A Comprehensive Survey on Open-domain Question Answering
- Deep Contextualized Acoustic Representations For Semi-Supervised Speech Recognition
- Regression Transformer: Concurrent sequence regression and generation for molecular language modeling
- Adapt or Get Left Behind: Domain Adaptation through BERT Language Model Finetuning for Aspect-Target Sentiment Classification
- ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning
- Neural Networks for Entity Matching: A Survey
- Measuring and Reducing Gendered Correlations in Pre-trained Models
- MetaSleepLearner: A Pilot Study on Fast Adaptation of Bio-signals-Based Sleep Stage Classifier to New Individual Subject Using Meta-Learning
- Improving Transformer-based Speech Recognition Using Unsupervised Pre-training
- A Survey on Text Classification: From Shallow to Deep Learning
- Mixout: Effective Regularization to Finetune Large-scale Pretrained Language Models
- StructBERT: Incorporating Language Structures into Pre-training for Deep Language Understanding
- Generating Accurate Assert Statements for Unit Test Cases using Pretrained Transformers
- Beyond 512 Tokens: Siamese Multi-depth Transformer-based Hierarchical Encoder for Long-Form Document Matching
- Pre-training via Paraphrasing
- Efficient Attention: Attention with Linear Complexities
- NEZHA: Neural Contextualized Representation for Chinese Language Understanding
- Low-Resource Knowledge-Grounded Dialogue Generation
- Language Models are Open Knowledge Graphs
- DepressionNet: A Novel Summarization Boosted Deep Framework for Depression Detection on Social Media
- Secure Evaluation of Quantized Neural Networks
- Structured Pruning of a BERT-based Question Answering Model
- Overview of the TREC 2019 deep learning track
- MisRoBÆRTa: Transformers versus Misinformation
- Rethinking Positional Encoding in Language Pre-training
- A Comparison of LSTM and BERT for Small Corpus
- Classification Aware Neural Topic Model and its Application on a New COVID-19 Disinformation Corpus
- MEANTIME: Mixture of Attention Mechanisms with Multi-temporal Embeddings for Sequential Recommendation
- Med-BERT: pre-trained contextualized embeddings on large-scale structured electronic health records for disease prediction
- Recent Advances in Natural Language Inference: A Survey of Benchmarks, Resources, and Approaches
- A Survey of Deep Learning for Scientific Discovery
- FastBERT: a Self-distilling BERT with Adaptive Inference Time
- Generalization through Memorization: Nearest Neighbor Language Models
- Gradient Boosting Neural Networks: GrowNet
- The Eighth Dialog System Technology Challenge
- MathBERT: A Pre-Trained Model for Mathematical Formula Understanding
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT
- NAS-BERT: Task-Agnostic and Adaptive-Size BERT Compression with Neural Architecture Search
- Zero-Resource Knowledge-Grounded Dialogue Generation
- CLUENER2020: Fine-grained Named Entity Recognition Dataset and Benchmark for Chinese
- Using word embeddings to improve the discriminability of co-occurrence text networks
- Cross-lingual Retrieval for Iterative Self-Supervised Training
- Reweighted Proximal Pruning for Large-Scale Language Representation
- Explaining Question Answering Models through Text Generation
- Are Transformers universal approximators of sequence-to-sequence functions?
- Using Prior Knowledge to Guide BERT's Attention in Semantic Textual Matching Tasks
- Technical report on Conversational Question Answering
- Neural ranking models for document retrieval
- Automated Quality Assessment of Cognitive Behavioral Therapy Sessions Through Highly Contextualized Language Representations
- FastMoE: A Fast Mixture-of-Expert Training System
- Recent progress in the JARVIS infrastructure for next-generation data-driven materials design
- Factorized Multimodal Transformer for Multimodal Sequential Learning
- Addressing "Documentation Debt" in Machine Learning Research: A Retrospective Datasheet for BookCorpus
- 12-in-1: Multi-Task Vision and Language Representation Learning
- Pre-trained Language Models in Biomedical Domain: A Systematic Survey
- Are Pre-trained Language Models Aware of Phrases? Simple but Strong Baselines for Grammar Induction
- Robustness Verification for Transformers
- What Would Elsa Do? Freezing Layers During Transformer Fine-Tuning
- The Use of a Large Language Model for Cyberbullying Detection
- LOREN: Logic-Regularized Reasoning for Interpretable Fact Verification
- Dice Loss for Data-imbalanced NLP Tasks
- IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP
- Corpus-Based Paraphrase Detection Experiments and Review
- Accelerating Training of Transformer-Based Language Models with Progressive Layer Dropping
- Rethinking Positional Encoding
- End-to-end Named Entity Recognition and Relation Extraction using Pre-trained Language Models
- Transformer-Based Models for Automatic Identification of Argument Relations: A Cross-Domain Evaluation
- Outside the Box: Abstraction-Based Monitoring of Neural Networks
- SKEP: Sentiment Knowledge Enhanced Pre-training for Sentiment Analysis
- Socially Enhanced Situation Awareness from Microblogs using Artificial Intelligence: A Survey
- TRANS-BLSTM: Transformer with Bidirectional LSTM for Language Understanding
- Tree-structured Attention with Hierarchical Accumulation
- PipeMare: Asynchronous Pipeline Parallel DNN Training
- Pretrained Encyclopedia: Weakly Supervised Knowledge-Pretrained Language Model
- Fake or Genuine? Contextualised Text Representation for Fake Review Detection
- Learning to Encode Position for Transformer with Continuous Dynamical Model
- Ultra-High-Resolution Detector Simulation with Intra-Event Aware GAN and Self-Supervised Relational Reasoning
- A Review on the Applications of Transformer-based language models for Nucleotide Sequence Analysis
- Reducing Sentiment Bias in Language Models via Counterfactual Evaluation
- Taking a Stance on Fake News: Towards Automatic Disinformation Assessment via Deep Bidirectional Transformer Language Models for Stance Detection
- A Comparative Study of Transformer-Based Language Models on Extractive Question Answering
- The Effect of Natural Distribution Shift on Question Answering Models
- OPT: Omni-Perception Pre-Trainer for Cross-Modal Understanding and Generation
- EVA2.0: Investigating Open-Domain Chinese Dialogue Systems with Large-Scale Pre-Training
- Graph-Based Reasoning over Heterogeneous External Knowledge for Commonsense Question Answering
- TreeBERT: A Tree-Based Pre-Trained Model for Programming Language
- Semantic Networks for Engineering Design: A Survey
- XGPT: Cross-modal Generative Pre-Training for Image Captioning
- Certified Data Removal from Machine Learning Models
- CORD19STS: COVID-19 Semantic Textual Similarity Dataset
- DeeBERT: Dynamic Early Exiting for Accelerating BERT Inference
- KELM: Knowledge Enhanced Pre-Trained Language Representations with Message Passing on Hierarchical Relational Graphs
- Pretrained AI Models: Performativity, Mobility, and Change
- SeMemNN: A Semantic Matrix-Based Memory Neural Network for Text Classification
- Vision Transformers: From Semantic Segmentation to Dense Prediction
- Automated Novelty Evaluation of Academic Paper: A Collaborative Approach Integrating Human and Large Language Model Knowledge
- Evolution of transfer learning in natural language processing
- Speaker-Aware BERT for Multi-Turn Response Selection in Retrieval-Based Chatbots
- Conformer-Kernel with Query Term Independence for Document Retrieval
- On Linear Identifiability of Learned Representations
- Fast Multi-Resolution Transformer Fine-tuning for Extreme Multi-label Text Classification
- Exploring Benefits of Transfer Learning in Neural Machine Translation
- Overview of the TREC 2022 deep learning track
- Unsupervised Domain Adaptation of Language Models for Reading Comprehension
- Pre-Trained Models: Past, Present and Future
- ProTo: Program-Guided Transformer for Program-Guided Tasks
- Diverse Image Inpainting with Bidirectional and Autoregressive Transformers
- SOLD: Sinhala Offensive Language Dataset
- DeepSeek-V3, GPT-4, Phi-4, and LLaMA-3.3 generate correct code for LoRaWAN-related engineering tasks
- Self-supervised Image-text Pre-training With Mixed Data In Chest X-rays
- Towards Domain Adaptation from Limited Data for Question Answering Using Deep Neural Networks
- EmotionIC: emotional inertia and contagion-driven dependency modeling for emotion recognition in conversation
- Look at What I'm Doing: Self-Supervised Spatial Grounding of Narrations in Instructional Videos
- "When they say weed causes depression, but it's your fav antidepressant": Knowledge-aware Attention Framework for Relationship Extraction
- Pre-Trained and Attention-Based Neural Networks for Building Noetic Task-Oriented Dialogue Systems
- QUEACO: Borrowing Treasures from Weakly-labeled Behavior Data for Query Attribute Value Extraction
- Cracking the Contextual Commonsense Code: Understanding Commonsense Reasoning Aptitude of Deep Contextual Representations
- Taming Pretrained Transformers for Extreme Multi-label Text Classification
- Improved Hierarchical Patient Classification with Language Model Pretraining over Clinical Notes
- LMSanitator: Defending Prompt-Tuning Against Task-Agnostic Backdoors
- Neuro-symbolic Architectures for Context Understanding
- DeepSI: Interactive Deep Learning for Semantic Interaction
- NLP From Scratch Without Large-Scale Pretraining: A Simple and Efficient Framework
- MC-BERT: Efficient Language Pre-Training via a Meta Controller
- Analyzing the Forgetting Problem in the Pretrain-Finetuning of Dialogue Response Models
- Transfer Learning from Transformers to Fake News Challenge Stance Detection (FNC-1) Task
- Improve Transformer Models with Better Relative Position Embeddings
- Spatial-Temporal Self-Attention Network for Flow Prediction
- On the comparability of Pre-trained Language Models
- Lacking the embedding of a word? Look it up into a traditional dictionary
- Multi-task Batch Reinforcement Learning with Metric Learning
- Enabling Language Models to Fill in the Blanks
- Learning Norms from Stories: A Prior for Value Aligned Agents
- Retrieval-Augmented Transformer-XL for Close-Domain Dialog Generation
- Unsupervised Editing for Counterfactual Stories
- Defending Against Backdoor Attacks in Natural Language Generation
- SCARF: Self-Supervised Contrastive Learning using Random Feature Corruption
- Training Multilingual Pre-trained Language Model with Byte-level Subwords
- Real-Time Execution of Large-scale Language Models on Mobile
- Connections are Expressive Enough: Universal Approximability of Sparse Transformers
- DocVQA: A Dataset for VQA on Document Images
- Variance-reduced Language Pretraining via a Mask Proposal Network
- VIOLIN: A Large-Scale Dataset for Video-and-Language Inference
- Multi-Stage Conversational Passage Retrieval: An Approach to Fusing Term Importance Estimation and Neural Query Rewriting
- Scatterbrain: Unifying Sparse and Low-rank Attention Approximation
- LightPAFF: A Two-Stage Distillation Framework for Pre-training and Fine-tuning
- Mining Implicit Relevance Feedback from User Behavior for Web Question Answering
- Literature review on vulnerability detection using NLP technology
- DeepTrax: Embedding Graphs of Financial Transactions
- Adversarial NLI for Factual Correctness in Text Summarisation Models
- Sentence Analogies: Exploring Linguistic Relationships and Regularities in Sentence Embeddings
- Vokenization: Improving Language Understanding with Contextualized, Visual-Grounded Supervision
- Denoised Internal Models: a Brain-Inspired Autoencoder against Adversarial Attacks
- Sketch-BERT: Learning Sketch Bidirectional Encoder Representation from Transformers by Self-supervised Learning of Sketch Gestalt
- Entity and Evidence Guided Relation Extraction for DocRED
- Fusing Context Into Knowledge Graph for Commonsense Question Answering
- Disentangling Adaptive Gradient Methods from Learning Rates
- CoTexT: Multi-task Learning with Code-Text Transformer
- Evaluation of Sentence Representations in Polish
- A Comparison of Pre-trained Vision-and-Language Models for Multimodal Representation Learning across Medical Images and Reports
- Multi-node Bert-pretraining: Cost-efficient Approach
- Efficient Softmax Approximation for Deep Neural Networks with Attention Mechanism
- InstructNet: A Novel Approach for Multi-Label Instruction Classification through Advanced Deep Learning
- Hide-and-Seek: A Template for Explainable AI
- Detecting and analyzing missing citations to published scientific entities
- Beyond English-Only Reading Comprehension: Experiments in Zero-Shot Multilingual Transfer for Bulgarian
- Question Generation by Transformers
- On the Impact of Knowledge-based Linguistic Annotations in the Quality of Scientific Embeddings
- E-Stitchup: Data Augmentation for Pre-Trained Embeddings
- An Analysis on Matching Mechanisms and Token Pruning for Late-interaction Models
- Symmetric Regularization based BERT for Pair-wise Semantic Reasoning
- NUIG-Shubhanker@Dravidian-CodeMix-FIRE2020: Sentiment Analysis of Code-Mixed Dravidian text using XLNet
- Using Transformers to Provide Teachers with Personalized Feedback on their Classroom Discourse: The TalkMoves Application
- How Can BERT Help Lexical Semantics Tasks?
- An Algorithm for Routing Capsules in All Domains
- DLGNet: A Transformer-based Model for Dialogue Response Generation
- Evolving Character-level Convolutional Neural Networks for Text Classification
- LRC-BERT: Latent-representation Contrastive Knowledge Distillation for Natural Language Understanding
- Generative Deep Learning Techniques for Password Generation
- TTTTTackling WinoGrande Schemas
- Applying Transfer Learning for Improving Domain-Specific Search Experience Using Query to Question Similarity
- Robustly Pre-trained Neural Model for Direct Temporal Relation Extraction
- The CoRa Tensor Compiler: Compilation for Ragged Tensors with Minimal Padding
- Efficient Nearest Neighbor Language Models
- FMMformer: Efficient and Flexible Transformer via Decomposed Near-field and Far-field Attention
- A Comparative Study of Lexical Substitution Approaches based on Neural Language Models
- Low-Shot Classification: A Comparison of Classical and Deep Transfer Machine Learning Approaches
- Natural Language Inference in Context -- Investigating Contextual Reasoning over Long Texts
- Syntax-driven Iterative Expansion Language Models for Controllable Text Generation
- Cluster-based Deep Ensemble Learning for Emotion Classification in Internet Memes
- Optimizing Deeper Transformers on Small Datasets
- Audio MFCC-gram Transformers for respiratory insufficiency detection in COVID-19
- memeBot: Towards Automatic Image Meme Generation
- Curb Your Carbon Emissions: Benchmarking Carbon Emissions in Machine Translation
- Position Masking for Language Models
- Explicit Pairwise Word Interaction Modeling Improves Pretrained Transformers for English Semantic Similarity Tasks
- Exploring Transformers in Emotion Recognition: a comparison of BERT, DistillBERT, RoBERTa, XLNet and ELECTRA
- PatentTransformer-2: Controlling Patent Text Generation by Structural Metadata
- Effective Transfer Learning for Identifying Similar Questions: Matching User Questions to COVID-19 FAQs
- MVP-BERT: Redesigning Vocabularies for Chinese BERT and Multi-Vocab Pretraining
- Invenio: Discovering Hidden Relationships Between Tasks/Domains Using Structured Meta Learning
- How to Select One Among All? An Extensive Empirical Study Towards the Robustness of Knowledge Distillation in Natural Language Understanding
- Unsupervised Pre-training for Natural Language Generation: A Literature Review
- Transfer Learning for Multi-lingual Tasks -- a Survey
- Does BERT Understand Sentiment? Leveraging Comparisons Between Contextual and Non-Contextual Embeddings to Improve Aspect-Based Sentiment Models
- Data-driven models and computational tools for neurolinguistics: a language technology perspective
- DEUX: An Attribute-Guided Framework for Sociable Recommendation Dialog Systems
- Optimizing small BERTs trained for German NER
- baller2vec++: A Look-Ahead Multi-Entity Transformer For Modeling Coordinated Agents
- CMV-BERT: Contrastive multi-vocab pretraining of BERT
- Understanding Semantics from Speech Through Pre-training
- Heads-up! Unsupervised Constituency Parsing via Self-Attention Heads
- A Survey of Quantum Theory Inspired Approaches to Information Retrieval
- Neural, Symbolic and Neural-Symbolic Reasoning on Knowledge Graphs
- Neural Composition: Learning to Generate from Multiple Models
- Transfer Learning Robustness in Multi-Class Categorization by Fine-Tuning Pre-Trained Contextualized Language Models
- Complex Transformer: A Framework for Modeling Complex-Valued Sequence
- Conversational Question Answering: A Survey
- Contextualized End-to-End Neural Entity Linking
- An Empirical Investigation of Contextualized Number Prediction
- A Qualitative Evaluation of Language Models on Automatic Question-Answering for COVID-19
- Masked ELMo: An evolution of ELMo towards fully contextual RNN language models
- Bilingual Language Modeling, A transfer learning technique for Roman Urdu
- Spoiler Alert: Using Natural Language Processing to Detect Spoilers in Book Reviews
- Current Limitations of Language Models: What You Need is Retrieval
- An Experimental Evaluation of Transformer-based Language Models in the Biomedical Domain
- Cocktail: Leveraging Ensemble Learning for Optimized Model Serving in Public Cloud
- Alignment Attention by Matching Key and Query Distributions
- Can You Fool AI by Doing a 180? $\unicode{x2013}$ A Case Study on Authorship Analysis of Texts by Arata Osada
- Quantifying the Task-Specific Information in Text-Based Classifications
- Improving Diversity of Neural Text Generation via Inverse Probability Weighting
- A Mixture of Heads is Better than Heads
- Billion-scale Pre-trained E-commerce Product Knowledge Graph Model
- An Effective Contextual Language Modeling Framework for Speech Summarization with Augmented Features
- A Pairwise Probe for Understanding BERT Fine-Tuning on Machine Reading Comprehension
- JASS: Japanese-specific Sequence to Sequence Pre-training for Neural Machine Translation
- Exploring Fluent Query Reformulations with Text-to-Text Transformers and Reinforcement Learning
- R&R: Metric-guided Adversarial Sentence Generation
- KoreALBERT: Pretraining a Lite BERT Model for Korean Language Understanding
- Fine-tuning Pre-trained Contextual Embeddings for Citation Content Analysis in Scholarly Publication
- Learning to Emphasize: Dataset and Shared Task Models for Selecting Emphasis in Presentation Slides
- A Deep Recurrent Survival Model for Unbiased Ranking
- Integrating Lexical Knowledge in Word Embeddings using Sprinkling and Retrofitting
- Pre-training for Ad-hoc Retrieval: Hyperlink is Also You Need
- Multilingual Medical Question Answering and Information Retrieval for Rural Health Intelligence Access
- An Overview on Generative AI at Scale with Edge-Cloud Computing
- WIKIR: A Python toolkit for building a large-scale Wikipedia-based English Information Retrieval Dataset
- Should We Be Pre-training? An Argument for End-task Aware Training as an Alternative
- Enhancing Review Comprehension with Domain-Specific Commonsense
- Finnish Language Modeling with Deep Transformer Models
- Applying Text Mining to Analyze Human Question Asking in Creativity Research
- Improving the quality of Persian clinical text with a novel spelling correction system
- PairConnect: A Compute-Efficient MLP Alternative to Attention
- Mischief: A Simple Black-Box Attack Against Transformer Architectures
- A Short Survey of Pre-trained Language Models for Conversational AI-A NewAge in NLP
- Larger-Context Tagging: When and Why Does It Work?
- Modeling Inter-Speaker Relationship in XLNet for Contextual Spoken Language Understanding
- On the Importance of Adaptive Data Collection for Extremely Imbalanced Pairwise Tasks
- ActBERT: Learning Global-Local Video-Text Representations
- Retraining DistilBERT for a Voice Shopping Assistant by Using Universal Dependencies
- TVDIM: Enhancing Image Self-Supervised Pretraining via Noisy Text Data
- Quantum-inspired Multimodal Fusion for Video Sentiment Analysis
- Multiplicative Position-aware Transformer Models for Language Understanding
- Improved Customer Transaction Classification using Semi-Supervised Knowledge Distillation
- Using Artificial Intelligence to Augment Science Prioritization for Astro2020
- Normalization of Input-output Shared Embeddings in Text Generation Models
- Transformer-based language modeling and decoding for conversational speech recognition
- Memeify: A Large-Scale Meme Generation System
- Multi-label Classification for Automatic Tag Prediction in the Context of Programming Challenges
- A neural document language modeling framework for spoken document retrieval
- Towards Learning Cross-Modal Perception-Trace Models
- BIG MOOD: Relating Transformers to Explicit Commonsense Knowledge
- Pre-train, Interact, Fine-tune: A Novel Interaction Representation for Text Classification
- Effective Unsupervised Domain Adaptation with Adversarially Trained Language Models
- Examining the rhetorical capacities of neural language models
- MaP: A Matrix-based Prediction Approach to Improve Span Extraction in Machine Reading Comprehension
- No Answer is Better Than Wrong Answer: A Reflection Model for Document Level Machine Reading Comprehension
- Using the Hammer Only on Nails: A Hybrid Method for Evidence Retrieval for Question Answering
- Emotion Correlation Mining Through Deep Learning Models on Natural Language Text
- DFraud3- Multi-Component Fraud Detection freeof Cold-start
- BERT Embeddings Can Track Context in Conversational Search
- Evaluating Document Coherence Modelling
- Learning a Word-Level Language Model with Sentence-Level Noise Contrastive Estimation for Contextual Sentence Probability Estimation
- LRG at TREC 2020: Document Ranking with XLNet-Based Models
- Few Shot Learning for Information Verification
- Improving speech recognition models with small samples for air traffic control systems
- Multiple Structural Priors Guided Self Attention Network for Language Understanding
- IIT_kgp at FinCausal 2020, Shared Task 1: Causality Detection using Sentence Embeddings in Financial Reports
- Character-level Representations Improve DRS-based Semantic Parsing Even in the Age of BERT
- Towards Preference Learning for Autonomous Ground Robot Navigation Tasks
- What makes us curious? analysis of a corpus of open-domain questions
- One-shot Key Information Extraction from Document with Deep Partial Graph Matching
- Small-Bench NLP: Benchmark for small single GPU trained models in Natural Language Processing
- Don't be Contradicted with Anything! CI-ToD: Towards Benchmarking Consistency for Task-oriented Dialogue System
- Unsupervised Keyphrase Extraction by Jointly Modeling Local and Global Context
- Semantic Answer Type Prediction using BERT: IAI at the ISWC SMART Task 2020
- Musical Speech: A Transformer-based Composition Tool
- Knowledge Transfer by Discriminative Pre-training for Academic Performance Prediction
- The Brownian motion in the transformer model
- LV-BERT: Exploiting Layer Variety for BERT
- Why Can You Lay Off Heads? Investigating How BERT Heads Transfer
- Leveraging Linguistic Coordination in Reranking N-Best Candidates For End-to-End Response Selection Using BERT
- Occlusion Aware Kernel Correlation Filter Tracker using RGB-D
- Towards a Transformer-Based Reverse Dictionary Model for Quality Estimation of Definitions
- Action based Network for Conversation Question Reformulation
- Answer Generation for Questions With Multiple Information Sources in E-Commerce
- An Empirical Study of Topic Transition in Dialogue
- Cross-language Information Retrieval
- When to Fold'em: How to answer Unanswerable questions
- Modelling General Properties of Nouns by Selectively Averaging Contextualised Embeddings
- Bi-directional Cognitive Thinking Network for Machine Reading Comprehension
- Pre-trained Language Model Based Active Learning for Sentence Matching
- Pre-Trained Models for Heterogeneous Information Networks
- Fine-tuning Multi-hop Question Answering with Hierarchical Graph Network
- Integrating Knowledge into End-to-End Speech Recognition from External Text-Only Data
- A Morpho-Syntactically Informed LSTM-CRF Model for Named Entity Recognition