Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
arXiv:1910.10683
Abstract
Transfer learning, where a model is first pre-trained on a data-rich task before being fine-tuned on a downstream task, has emerged as a powerful technique in natural language processing (NLP). The effectiveness of transfer learning has given rise to a diversity of approaches, methodology, and practice. In this paper, we explore the landscape of transfer learning techniques for NLP by introducing a unified framework that converts all text-based language problems into a text-to-text format. Our systematic study compares pre-training objectives, architectures, unlabeled data sets, transfer approaches, and other factors on dozens of language understanding tasks. By combining the insights from our exploration with scale and our new ``Colossal Clean Crawled Corpus'', we achieve state-of-the-art results on many benchmarks covering summarization, question answering, text classification, and more. To facilitate future work on transfer learning for NLP, we release our data set, pre-trained models, and code.
References in corpus (22)
- Distilling the Knowledge in a Neural Network
- Sequence to Sequence Learning with Neural Networks
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- How transferable are features in deep neural networks?
- An Overview of Multi-Task Learning in Deep Neural Networks
- Cross-lingual Language Model Pretraining
- ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
- One weird trick for parallelizing convolutional neural networks
- CTRL: A Conditional Transformer Language Model for Controllable Generation
- Skip-Thought Vectors
- MASS: Masked Sequence to Sequence Pre-training for Language Generation
- Deep Learning Scaling is Predictable, Empirically
- What makes ImageNet good for transfer learning?
- Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges
- BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
- Parameter-Efficient Transfer Learning for NLP
- TinyBERT: Distilling BERT for Natural Language Understanding
- StructBERT: Incorporating Language Structures into Pre-training for Deep Language Understanding
- NewsQA: A Machine Comprehension Dataset
- Generating Wikipedia by Summarizing Long Sequences
- The Bottom-up Evolution of Representations in the Transformer: A Study with Machine Translation and Language Modeling Objectives
- SummAE: Zero-Shot Abstractive Text Summarization using Length-Agnostic Auto-Encoders
Cited by in corpus (575)
- Learning Transferable Visual Models From Natural Language Supervision
- Survey of Hallucination in Natural Language Generation
- Transformers in Vision: A Survey
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence
- Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing
- Pre-trained Models for Natural Language Processing: A Survey
- Is Space-Time Attention All You Need for Video Understanding?
- ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
- Linformer: Self-Attention with Linear Complexity
- Multi-Task Learning for Dense Prediction Tasks: A Survey
- Adapted Large Language Models Can Outperform Medical Experts in Clinical Text Summarization
- CoAtNet: Marrying Convolution and Attention for All Data Sizes
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers
- Multilingual Denoising Pre-training for Neural Machine Translation
- CORD-19: The COVID-19 Open Research Dataset
- MPNet: Masked and Permuted Pre-training for Language Understanding
- REALM: Retrieval-Augmented Language Model Pre-Training
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Big Self-Supervised Models are Strong Semi-Supervised Learners
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- ChatGPT for Shaping the Future of Dentistry: The Potential of Multi-Modal Large Language Model
- Big Bird: Transformers for Longer Sequences
- P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks
- CodeBERT: A Pre-Trained Model for Programming and Natural Languages
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training
- Should ChatGPT be Biased? Challenges and Risks of Bias in Large Language Models
- ASE: Large-Scale Reusable Adversarial Skill Embeddings for Physically Simulated Characters
- Controllable Protein Design with Language Models
- WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning
- Multi-Stage Document Ranking with BERT
- Post-hoc Interpretability for Neural NLP: A Survey
- Long Range Arena: A Benchmark for Efficient Transformers
- What is being transferred in transfer learning?
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
- Listen, Denoise, Action! Audio-Driven Motion Synthesis with Diffusion Models
- Portuguese Named Entity Recognition using BERT-CRF
- Transformers in Healthcare: A Survey
- GraphCodeBERT: Pre-training Code Representations with Data Flow
- Towards Unified Conversational Recommender Systems via Knowledge-Enhanced Prompt Learning
- How to DP-fy ML: A Practical Guide to Machine Learning with Differential Privacy
- Neural Program Repair with Execution-based Backpropagation
- Retrieving and Reading: A Comprehensive Survey on Open-domain Question Answering
- A Survey of Natural Language Generation
- Time-Aware Language Models as Temporal Knowledge Bases
- Self-Supervised Learning for Videos: A Survey
- Regression Transformer: Concurrent sequence regression and generation for molecular language modeling
- Toward Transformer-Based Object Detection
- Data Augmentation using Pre-trained Transformer Models
- Self-supervised Pretraining of Visual Features in the Wild
- K-Adapter: Infusing Knowledge into Pre-Trained Models with Adapters
- TinyBERT: Distilling BERT for Natural Language Understanding
- GPT-3-driven pedagogical agents for training children's curious question-asking skills
- Multilingual is not enough: BERT for Finnish
- Semantic Models for the First-stage Retrieval: A Comprehensive Review
- Biomedical Question Answering: A Survey of Approaches and Challenges
- OctFormer: Octree-based Transformers for 3D Point Clouds
- An Empirical Survey on Long Document Summarization: Datasets, Models and Metrics
- CPT: A Pre-Trained Unbalanced Transformer for Both Chinese Language Understanding and Generation
- Measuring and Reducing Gendered Correlations in Pre-trained Models
- WT5?! Training Text-to-Text Models to Explain their Predictions
- Pretrained Transformers as Universal Computation Engines
- Generative Data Augmentation for Commonsense Reasoning
- An Empirical Study on the Usage of Transformer Models for Code Completion
- Generating Accurate Assert Statements for Unit Test Cases using Pretrained Transformers
- RiNALMo: General-Purpose RNA Language Models Can Generalize Well on Structure Prediction Tasks
- Rethinking Search: Making Domain Experts out of Dilettantes
- LLM for SoC Security: A Paradigm Shift
- Pre-training via Paraphrasing
- Paradigm Shift in Natural Language Processing
- Measuring the Algorithmic Efficiency of Neural Networks
- CPT: Colorful Prompt Tuning for Pre-trained Vision-Language Models
- GREEK-BERT: The Greeks visiting Sesame Street
- ProphetNet: Predicting Future N-gram for Sequence-to-Sequence Pre-training
- WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training
- Multi-task learning for natural language processing in the 2020s: where are we going?
- Structured Denoising Diffusion Models in Discrete State-Spaces
- SelfAPR: Self-supervised Program Repair with Test Execution Diagnostics
- Compositional Generalization in Semantic Parsing: Pre-training vs. Specialized Architectures
- Zero-shot Text Classification With Generative Language Models
- Few-Shot Conversational Dense Retrieval
- Structured Pruning of a BERT-based Question Answering Model
- A Survey of Large Language Models for Graphs
- Full Page Handwriting Recognition via Image to Sequence Extraction
- Recursively Summarizing Books with Human Feedback
- Rethinking Positional Encoding in Language Pre-training
- Sabiá: Portuguese Large Language Models
- Augmenting Low-Resource Text Classification with Graph-Grounded Pre-training and Prompting
- BASE Layers: Simplifying Training of Large, Sparse Models
- Rapidly Bootstrapping a Question Answering Dataset for COVID-19
- Classification of Human- and AI-Generated Texts: Investigating Features for ChatGPT
- RepBERT: Contextualized Text Embeddings for First-Stage Retrieval
- Bidirectional Generation of Structure and Properties Through a Single Molecular Foundation Model
- JARVIS-Leaderboard: A Large Scale Benchmark of Materials Design Methods
- Augmenting Interpretable Models with LLMs during Training
- What does BERT know about books, movies and music? Probing BERT for Conversational Recommendation
- FastBERT: a Self-distilling BERT with Adaptive Inference Time
- Space-time Mixing Attention for Video Transformer
- One-Shot Labeling for Automatic Relevance Estimation
- CorpusBrain: Pre-train a Generative Retrieval Model for Knowledge-Intensive Language Tasks
- A Survey on Transfer Learning in Natural Language Processing
- Detecting AI-Generated Text: Factors Influencing Detectability with Current Methods
- HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
- Improving accuracy of GPT-3/4 results on biomedical data using a retrieval-augmented language model
- RussianSuperGLUE: A Russian Language Understanding Evaluation Benchmark
- Explaining the Explainers in Graph Neural Networks: a Comparative Study
- Cross-lingual Retrieval for Iterative Self-Supervised Training
- Contextual Feature Extraction Hierarchies Converge in Large Language Models and the Brain
- I-MAD: Interpretable Malware Detector Using Galaxy Transformer
- Generating Bug-Fixes Using Pretrained Transformers
- Self-training Improves Pre-training for Natural Language Understanding
- Towards Zero-Label Language Learning
- Explainable Legal Case Matching via Inverse Optimal Transport-based Rationale Extraction
- Explaining Question Answering Models through Text Generation
- BERT Loses Patience: Fast and Robust Inference with Early Exit
- Knowledge-Aware Language Model Pretraining
- Generative Language Modeling for Automated Theorem Proving
- Mask and Reason: Pre-Training Knowledge Graph Transformers for Complex Logical Queries
- Automated Data Visualization from Natural Language via Large Language Models: An Exploratory Study
- Creativity and Machine Learning: A Survey
- Learning the Travelling Salesperson Problem Requires Rethinking Generalization
- Alignment of Language Agents
- Distilling Knowledge from Reader to Retriever for Question Answering
- Notable: On-the-fly Assistant for Data Storytelling in Computational Notebooks
- How Can We Know When Language Models Know? On the Calibration of Language Models for Question Answering
- Let Me Do It For You: Towards LLM Empowered Recommendation via Tool Learning
- Towards Continual Knowledge Learning of Language Models
- Vulnerabilities in AI Code Generators: Exploring Targeted Data Poisoning Attacks
- Recent progress in the JARVIS infrastructure for next-generation data-driven materials design
- Diet Code Is Healthy: Simplifying Programs for Pre-trained Models of Code
- Semantic Gesticulator: Semantics-Aware Co-Speech Gesture Synthesis
- AMMUS : A Survey of Transformer-based Pretrained Models in Natural Language Processing
- FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device Placement
- Continual Learning for Generative Retrieval over Dynamic Corpora
- Learning to summarize from human feedback
- A Unified Review of Deep Learning for Automated Medical Coding
- Overview of BioASQ 2020: The eighth BioASQ challenge on Large-Scale Biomedical Semantic Indexing and Question Answering
- LipLearner: Customizable Silent Speech Interactions on Mobile Devices
- Local Interpretations for Explainable Natural Language Processing: A Survey
- Support-set bottlenecks for video-text representation learning
- Language Models as Few-Shot Learner for Task-Oriented Dialogue Systems
- Conversational Question Reformulation via Sequence-to-Sequence Architectures and Pretrained Language Models
- 12-in-1: Multi-Task Vision and Language Representation Learning
- Measuring a hate speech spectrum with faceted Rasch item response theory and perspective-aware, explainable-by-design deep learning
- Recipes for Adapting Pre-trained Monolingual and Multilingual Models to Machine Translation
- I Know What You Are Searching For: Code Snippet Recommendation from Stack Overflow Posts
- Rapidly Deploying a Neural Search Engine for the COVID-19 Open Research Dataset: Preliminary Thoughts and Lessons Learned
- Leveraging Monolingual Data with Self-Supervision for Multilingual Neural Machine Translation
- Towards Explainable Conversational Recommender Systems
- Graph Neural Networks in Vision-Language Image Understanding: A Survey
- LFPT5: A Unified Framework for Lifelong Few-shot Language Learning Based on Prompt Tuning of T5
- CLUECorpus2020: A Large-scale Chinese Corpus for Pre-training Language Model
- VIMPAC: Video Pre-Training via Masked Token Prediction and Contrastive Learning
- Are Pre-trained Convolutions Better than Pre-trained Transformers?
- AQuaMuSe: Automatically Generating Datasets for Query-Based Multi-Document Summarization
- GalleryGPT: Analyzing Paintings with Large Multimodal Models
- Unifying Two-Stream Encoders with Transformers for Cross-Modal Retrieval
- Downstream Model Design of Pre-trained Language Model for Relation Extraction Task
- Establishing Baselines for Text Classification in Low-Resource Languages
- HTLM: Hyper-Text Pre-Training and Prompting of Language Models
- GypSum: Learning Hybrid Representations for Code Summarization
- Adaptive Re-Ranking with a Corpus Graph
- CO-Search: COVID-19 Information Retrieval with Semantic Search, Question Answering, and Abstractive Summarization
- Out-of-Domain Semantics to the Rescue! Zero-Shot Hybrid Retrieval Models
- IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP
- PeCo: Perceptual Codebook for BERT Pre-training of Vision Transformers
- Few-shot Natural Language Generation for Task-Oriented Dialog
- Description Based Text Classification with Reinforcement Learning
- A Unified Generative Retriever for Knowledge-Intensive Language Tasks via Prompt Learning
- Clinical Insights: A Comprehensive Review of Language Models in Medicine
- Grounding and Evaluation for Large Language Models: Practical Challenges and Lessons Learned (Survey)
- CatAlyst: Domain-Extensible Intervention for Preventing Task Procrastination Using Large Generative Models
- Rethinking Positional Encoding
- Automatic Personalized Impression Generation for PET Reports Using Large Language Models
- A Paragraph-level Multi-task Learning Model for Scientific Fact-Verification
- A Memory Efficient Baseline for Open Domain Question Answering
- An empirical study of LLaMA3 quantization: from LLMs to MLLMs
- Example-Based Named Entity Recognition
- CodeEditor: Learning to Edit Source Code with Pre-trained Models
- ETC: Encoding Long and Structured Inputs in Transformers
- CoDA: Contrast-enhanced and Diversity-promoting Data Augmentation for Natural Language Understanding
- Combiner: Full Attention Transformer with Sparse Computation Cost
- ERNIE-GEN: An Enhanced Multi-Flow Pre-training and Fine-tuning Framework for Natural Language Generation
- SKEP: Sentiment Knowledge Enhanced Pre-training for Sentiment Analysis
- WikiBERT models: deep transfer learning for many languages
- On Feature Normalization and Data Augmentation
- Socially Enhanced Situation Awareness from Microblogs using Artificial Intelligence: A Survey
- Neural Program Repair: Systems, Challenges and Solutions
- Scalable Multi-Hop Relational Reasoning for Knowledge-Aware Question Answering
- TRANS-BLSTM: Transformer with Bidirectional LSTM for Language Understanding
- A Text-guided Protein Design Framework
- Applying Transformer-based Text Summarization for Keyphrase Generation
- Extracting Cultural Commonsense Knowledge at Scale
- Knowledge Enhanced Pretrained Language Models: A Compreshensive Survey
- Generative AI for Pull Request Descriptions: Adoption, Impact, and Developer Interventions
- Killing Two Birds with One Stone: Malicious Package Detection in NPM and PyPI using a Single Model of Malicious Behavior Sequence
- Interpreting Deep Learning Models in Natural Language Processing: A Review
- QuestEval: Summarization Asks for Fact-based Evaluation
- Entity-aware Transformers for Entity Search
- Learning to Encode Position for Transformer with Continuous Dynamical Model
- Comparison between parameter-efficient techniques and full fine-tuning: A case study on multilingual news article classification
- DocBank: A Benchmark Dataset for Document Layout Analysis
- Neural Language Generation: Formulation, Methods, and Evaluation
- Leveraging Large Language Models for Topic Classification in the Domain of Public Affairs
- The Dawn of AI-Native EDA: Opportunities and Challenges of Large Circuit Models
- EVA2.0: Investigating Open-Domain Chinese Dialogue Systems with Large-Scale Pre-Training
- GMAT: Global Memory Augmentation for Transformers
- Rankers, Judges, and Assistants: Towards Understanding the Interplay of LLMs in Information Retrieval Evaluation
- TrajGPT: Controlled Synthetic Trajectory Generation Using a Multitask Transformer-Based Spatiotemporal Model
- On Grounded Planning for Embodied Tasks with Language Models
- GLU Variants Improve Transformer
- Discourse-Aware Neural Extractive Text Summarization
- SemVLP: Vision-Language Pre-training by Aligning Semantics at Multiple Levels
- Portrayal: Leveraging NLP and Visualization for Analyzing Fictional Characters
- A Review of Winograd Schema Challenge Datasets and Approaches
- PALM: Pre-training an Autoencoding&Autoregressive Language Model for Context-conditioned Generation
- Learning Models of Individual Behavior in Chess
- Application of Deep Learning in Generating Structured Radiology Reports: A Transformer-Based Technique
- More but Correct: Generating Diversified and Entity-revised Medical Response
- Improving Readability for Automatic Speech Recognition Transcription
- Prior Knowledge Integration via LLM Encoding and Pseudo Event Regulation for Video Moment Retrieval
- KELM: Knowledge Enhanced Pre-Trained Language Representations with Message Passing on Hierarchical Relational Graphs
- Evaluating Biased Attitude Associations of Language Models in an Intersectional Context
- Reducing Energy Bloat in Large Model Training
- Document Ranking with a Pretrained Sequence-to-Sequence Model
- EarlyBERT: Efficient BERT Training via Early-bird Lottery Tickets
- Answering Complex Open-Domain Questions with Multi-Hop Dense Retrieval
- Enhancing User Behavior Sequence Modeling by Generative Tasks for Session Search
- End-to-End QA on COVID-19: Domain Adaptation with Synthetic Training
- Classification of Edge-dependent Labels of Nodes in Hypergraphs
- C5T5: Controllable Generation of Organic Molecules with Transformers
- A Statistical Framework of Watermarks for Large Language Models: Pivot, Detection Efficiency and Optimal Rules
- Multimodal Data Augmentation for Image Captioning using Diffusion Models
- Utterance-level Dialogue Understanding: An Empirical Study
- Generating Usage-related Questions for Preference Elicitation in Conversational Recommender Systems
- PyTy: Repairing Static Type Errors in Python
- I-BERT: Integer-only BERT Quantization
- Transferability of Natural Language Inference to Biomedical Question Answering
- Beyond Leaderboards: A survey of methods for revealing weaknesses in Natural Language Inference data and models
- The State of the Art in Creating Visualization Corpora for Automated Chart Analysis
- Stealing the Decoding Algorithms of Language Models
- WebCode2M: A Real-World Dataset for Code Generation from Webpage Designs
- EAGER-LLM: Enhancing Large Language Models as Recommenders through Exogenous Behavior-Semantic Integration
- Brain2Word: Decoding Brain Activity for Language Generation
- NewsEmbed: Modeling News through Pre-trained Document Representations
- A survey of textual cyber abuse detection using cutting-edge language models and large language models
- Explaining Bayesian Neural Networks
- Match-Prompt: Improving Multi-task Generalization Ability for Neural Text Matching via Prompt Learning
- MiniLMv2: Multi-Head Self-Attention Relation Distillation for Compressing Pretrained Transformers
- e-CLIP: Large-Scale Vision-Language Representation Learning in E-commerce
- PMI-Masking: Principled masking of correlated spans
- P^3 Ranker: Mitigating the Gaps between Pre-training and Ranking Fine-tuning with Prompt-based Learning and Pre-finetuning
- Efficient (Soft) Q-Learning for Text Generation with Limited Good Data
- The Microsoft Toolkit of Multi-Task Deep Neural Networks for Natural Language Understanding
- SpaDE: Improving Sparse Representations using a Dual Document Encoder for First-stage Retrieval
- Content-Based Collaborative Generation for Recommender Systems
- Constraint-aware and Ranking-distilled Token Pruning for Efficient Transformer Inference
- DeepDebug: Fixing Python Bugs Using Stack Traces, Backtranslation, and Code Skeletons
- ChartSumm: A Comprehensive Benchmark for Automatic Chart Summarization of Long and Short Summaries
- Emformer: Efficient Memory Transformer Based Acoustic Model For Low Latency Streaming Speech Recognition
- The Lean Data Scientist: Recent Advances towards Overcoming the Data Bottleneck
- PharmMT: A Neural Machine Translation Approach to Simplify Prescription Directions
- CausalBERT: Injecting Causal Knowledge Into Pre-trained Models with Minimal Supervision
- GenQREnsemble: Zero-Shot LLM Ensemble Prompting for Generative Query Reformulation
- Pre-training Text-to-Text Transformers for Concept-centric Common Sense
- LLM-Mediated Domain-Specific Voice Agents: The Case of TextileBot
- GUing: A Mobile GUI Search Engine using a Vision-Language Model
- Improving Sequential Recommendations with LLMs
- Syntactic Data Augmentation Increases Robustness to Inference Heuristics
- The Life Cycle of Knowledge in Big Language Models: A Survey
- SEAL: Segment-wise Extractive-Abstractive Long-form Text Summarization
- MinTL: Minimalist Transfer Learning for Task-Oriented Dialogue Systems
- The Language Interpretability Tool: Extensible, Interactive Visualizations and Analysis for NLP Models
- Hardware Accelerator for Multi-Head Attention and Position-Wise Feed-Forward in the Transformer
- Connecting the Dots: A Knowledgeable Path Generator for Commonsense Question Answering
- MORepair: Teaching LLMs to Repair Code via Multi-Objective Fine-tuning
- Review-driven Personalized Preference Reasoning with Large Language Models for Recommendation
- Deep Extrapolation for Attribute-Enhanced Generation
- Towards Question Format Independent Numerical Reasoning: A Set of Prerequisite Tasks
- Controlling Computation versus Quality for Neural Sequence Models
- Parallelizing Legendre Memory Unit Training
- Tiny Transformers for Environmental Sound Classification at the Edge
- Linguistically-driven Multi-task Pre-training for Low-resource Neural Machine Translation
- Improve Transformer Models with Better Relative Position Embeddings
- s2s-ft: Fine-Tuning Pretrained Transformer Encoders for Sequence-to-Sequence Learning
- TLDR: Token Loss Dynamic Reweighting for Reducing Repetitive Utterance Generation
- NLP From Scratch Without Large-Scale Pretraining: A Simple and Efficient Framework
- Word2Vec: Optimal Hyper-Parameters and Their Impact on NLP Downstream Tasks
- iScore: Visual Analytics for Interpreting How Language Models Automatically Score Summaries
- AdaptSum: Towards Low-Resource Domain Adaptation for Abstractive Summarization
- Language Models as Knowledge Bases: On Entity Representations, Storage Capacity, and Paraphrased Queries
- MC-BERT: Efficient Language Pre-Training via a Meta Controller
- Promptable Game Models: Text-Guided Game Simulation via Masked Diffusion Models
- Normalized Attention Without Probability Cage
- On Learning Universal Representations Across Languages
- POINTER: Constrained Progressive Text Generation via Insertion-based Generative Pre-training
- Enabling Language Models to Fill in the Blanks
- The Good, the Bad, and the Missing: Neural Code Generation for Machine Learning Tasks
- Zero-Shot Prompting and Few-Shot Fine-Tuning: Revisiting Document Image Classification Using Large Language Models
- INTERN: A New Learning Paradigm Towards General Vision
- Unveiling the Vulnerability of Private Fine-Tuning in Split-Based Frameworks for Large Language Models: A Bidirectionally Enhanced Attack
- Evolving Attention with Residual Convolutions
- Defending Against Backdoor Attacks in Natural Language Generation
- Deanthropomorphising NLP: Can a Language Model Be Conscious?
- Enhancing Multi-modal and Multi-hop Question Answering via Structured Knowledge and Unified Retrieval-Generation
- DynE: Dynamic Ensemble Decoding for Multi-Document Summarization
- Retrieval Augmented Zero-Shot Text Classification
- PTT5: Pretraining and validating the T5 model on Brazilian Portuguese data
- CBAG: Conditional Biomedical Abstract Generation
- A Few-shot Approach to Resume Information Extraction via Prompts
- Speech-language Pre-training for End-to-end Spoken Language Understanding
- To Tune or Not To Tune? Zero-shot Models for Legal Case Entailment
- DynaPipe: Optimizing Multi-task Training through Dynamic Pipelines
- LazyFormer: Self Attention with Lazy Update
- Advancing Multi-grained Alignment for Contrastive Language-Audio Pre-training
- Implicit Feedback for Dense Passage Retrieval: A Counterfactual Approach
- Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML Serving
- Mixed-Lingual Pre-training for Cross-lingual Summarization
- Length-Adaptive Transformer: Train Once with Length Drop, Use Anytime with Search
- CAiRE-COVID: A Question Answering and Query-focused Multi-Document Summarization System for COVID-19 Scholarly Information Management
- Auditing Gender Presentation Differences in Text-to-Image Models
- Evaluating representations by the complexity of learning low-loss predictors
- Question rewriting? Assessing its importance for conversational question answering
- Czert -- Czech BERT-like Model for Language Representation
- VX2TEXT: End-to-End Learning of Video-Based Text Generation From Multimodal Inputs
- Using Transformer based Ensemble Learning to classify Scientific Articles
- Amazon SageMaker Model Parallelism: A General and Flexible Framework for Large Model Training
- Unit Test Case Generation with Transformers and Focal Context
- A Hierarchical Network for Abstractive Meeting Summarization with Cross-Domain Pretraining
- Neural Passage Quality Estimation for Static Pruning
- Beyond Single Items: Exploring User Preferences in Item Sets with the Conversational Playlist Curation Dataset
- Vokenization: Improving Language Understanding with Contextualized, Visual-Grounded Supervision
- Do Language Models Understand Time?
- Fusing Context Into Knowledge Graph for Commonsense Question Answering
- Mirror: A Natural Language Interface for Data Querying, Summarization, and Visualization
- Reasoning Over Virtual Knowledge Bases With Open Predicate Relations
- An Enhanced Knowledge Injection Model for Commonsense Generation
- Progressive Entity Resolution: A Design Space Exploration
- Calibrated Language Model Fine-Tuning for In- and Out-of-Distribution Data
- Quam: Adaptive Retrieval through Query Affinity Modelling
- A Unified Transformer-based Framework for Duplex Text Normalization
- All-in-One Image-Grounded Conversational Agents
- ChroniclingAmericaQA: A Large-scale Question Answering Dataset based on Historical American Newspaper Pages
- Multi-View Sequence-to-Sequence Models with Conversational Structure for Abstractive Dialogue Summarization
- LAMBERT: Layout-Aware (Language) Modeling for information extraction
- MC-NN: An End-to-End Multi-Channel Neural Network Approach for Predicting Influenza A Virus Hosts and Antigenic Types
- Automatic Generation of Conversational Interfaces for Tabular Data Analysis
- Answering Questions on COVID-19 in Real-Time
- FedNLP: An interpretable NLP System to Decode Federal Reserve Communications
- Automatic Controllable Product Copywriting for E-Commerce
- StructFormer: Joint Unsupervised Induction of Dependency and Constituency Structure from Masked Language Modeling
- An Analysis on Matching Mechanisms and Token Pruning for Late-interaction Models
- SpeechNet: A Universal Modularized Model for Speech Processing Tasks
- Say What? Collaborative Pop Lyric Generation Using Multitask Transfer Learning
- A Survey of Spatio-Temporal EEG data Analysis: from Models to Applications
- CASPR: A Commonsense Reasoning-based Conversational Socialbot
- Lexically-constrained Text Generation through Commonsense Knowledge Extraction and Injection
- FAQ-Gen: An automated system to generate domain-specific FAQs to aid content comprehension
- English Prompts are Better for NLI-based Zero-Shot Emotion Classification than Target-Language Prompts
- DaCy: A Unified Framework for Danish NLP
- Anti-Distillation: Improving reproducibility of deep networks
- FathomGPT: A Natural Language Interface for Interactively Exploring Ocean Science Data
- Personalized Video Summarization by Multimodal Video Understanding
- Improving the Robustness of Transformer-based Large Language Models with Dynamic Attention
- ColdGANs: Taming Language GANs with Cautious Sampling Strategies
- Context-Enhanced Language Models for Generating Multi-Paper Citations
- Detection of Prosodic Boundaries in Speech Using Wav2Vec 2.0
- Causal Question Answering with Reinforcement Learning
- LazyTensor: combining eager execution with domain-specific compilers
- A Stack-Propagation Framework for Low-Resource Personalized Dialogue Generation
- Open-Domain Frame Semantic Parsing Using Transformers
- Towards Enriched Controllability for Educational Question Generation
- Which *BERT? A Survey Organizing Contextualized Encoders
- Galileo at SemEval-2020 Task 12: Multi-lingual Learning for Offensive Language Identification using Pre-trained Language Models
- Segmented Graph-Bert for Graph Instance Modeling
- Style Attuned Pre-training and Parameter Efficient Fine-tuning for Spoken Language Understanding
- CTRLStruct: Dialogue Structure Learning for Open-Domain Response Generation
- KGPT: Knowledge-Grounded Pre-Training for Data-to-Text Generation
- Transformer Based Language Models for Similar Text Retrieval and Ranking
- Reproducing NevIR: Negation in Neural Information Retrieval
- A Reproducibility Study of Goldilocks: Just-Right Tuning of BERT for TAR
- Linear Transformations for Cross-lingual Sentiment Analysis
- Folden: -Fold Ensemble for Out-Of-Distribution Detection
- Language Modeling using LMUs: 10x Better Data Efficiency or Improved Scaling Compared to Transformers
- Adaptive Name Entity Recognition under Highly Unbalanced Data
- A Modern Perspective on Query Likelihood with Deep Generative Retrieval Models
- Auto-MAP: A DQN Framework for Exploring Distributed Execution Plans for DNN Workloads
- Knowledge-Grounded Dialogue Generation with Pre-trained Language Models
- Direction is what you need: Improving Word Embedding Compression in Large Language Models
- HyperGrid: Efficient Multi-Task Transformers with Grid-wise Decomposable Hyper Projections
- Self-supervised pre-training and contrastive representation learning for multiple-choice video QA
- Scalable Multitask Learning Using Gradient-based Estimation of Task Affinity
- Zero-shot Generalization in Dialog State Tracking through Generative Question Answering
- Key Algorithms for Keyphrase Generation: Instruction-Based LLMs for Russian Scientific Keyphrases
- An Empirical Study on Neural Keyphrase Generation
- Compressed Deep Networks: Goodbye SVD, Hello Robust Low-Rank Approximation
- Train and You'll Miss It: Interactive Model Iteration with Weak Supervision and Pre-Trained Embeddings
- Natural Language Inference in Context -- Investigating Contextual Reasoning over Long Texts
- SeqGenSQL -- A Robust Sequence Generation Model for Structured Query Language
- Exemplar Guided Active Learning
- Fast Interleaved Bidirectional Sequence Generation
- RankingMatch: Delving into Semi-Supervised Learning with Consistency Regularization and Ranking Loss
- Neural Rule-Execution Tracking Machine For Transformer-Based Text Generation
- GlyphCRM: Bidirectional Encoder Representation for Chinese Character with its Glyph
- On the Effects of Regional Spelling Conventions in Retrieval Models
- A Canary in the AI Coal Mine: American Jews May Be Disproportionately Harmed by Intellectual Property Dispossession in Large Language Model Training
- SEPTQ: A Simple and Effective Post-Training Quantization Paradigm for Large Language Models
- Data-Driven Evolution of Library and Information Science Research Methods (1990-2022): A Perspective Based on Fine-grained Method Entities
- FairytaleQA Translated: Enabling Educational Question and Answer Generation in Less-Resourced Languages
- Predicting drug-gene relations via analogy tasks with word embeddings
- From Large to Tiny: Distilling and Refining Mathematical Expertise for Math Word Problems with Weakly Supervision
- Automated Description Generation of Cytologic Findings for Lung Cytological Images Using a Pretrained Vision Model and Dual Text Decoders: Preliminary Study
- Enhancing Source Code Classification Effectiveness via Prompt Learning Incorporating Knowledge Features
- Sentence-to-Label Generation Framework for Multi-task Learning of Japanese Sentence Classification and Named Entity Recognition
- M3PT: A Multi-Modal Model for POI Tagging
- AW-Opt: Learning Robotic Skills with Imitation and Reinforcement at Scale
- MT3: Multi-Task Multitrack Music Transcription
- 8-bit Optimizers via Block-wise Quantization
- EASE: Extractive-Abstractive Summarization with Explanations
- Are Larger Pretrained Language Models Uniformly Better? Comparing Performance at the Instance Level
- Testing pre-trained Transformer models for Lithuanian news clustering
- Exploring Versatile Generative Language Model Via Parameter-Efficient Transfer Learning
- Pre-training for Abstractive Document Summarization by Reinstating Source Text
- Giving Up Control: Neurons as Reinforcement Learning Agents
- Style is NOT a single variable: Case Studies for Cross-Style Language Understanding
- Lightweight and Direct Document Relevance Optimization for Generative Information Retrieval
- Is it feasible to detect FLOSS version release events from textual messages? A case study on Stack Overflow
- Go Wider Instead of Deeper
- Improving Commonsense Question Answering by Graph-based Iterative Retrieval over Multiple Knowledge Sources
- BEAMetrics: A Benchmark for Language Generation Evaluation Evaluation
- Automatic Summarization of Open-Domain Podcast Episodes
- Improving Self-supervised Pre-training via a Fully-Explored Masked Language Model
- Look at the First Sentence: Position Bias in Question Answering
- Source-Free Domain Adaptation for Question Answering with Masked Self-training
- PoinT-5: Pointer Network and T-5 based Financial NarrativeSummarisation
- To Pretrain or Not to Pretrain: Examining the Benefits of Pretraining on Resource Rich Tasks
- Constrained Auto-Regressive Decoding Constrains Generative Retrieval
- Investigating Task Arithmetic for Zero-Shot Information Retrieval
- Team Alex at CLEF CheckThat! 2020: Identifying Check-Worthy Tweets With Transformer Models
- Distilling Knowledge from Pre-trained Language Models via Text Smoothing
- EAGLE: Egocentric AGgregated Language-video Engine
- Improving Conversational Passage Re-ranking with View Ensemble
- FlanEC: Exploring Flan-T5 for Post-ASR Error Correction
- Harmful Terms and Where to Find Them: Measuring and Modeling Unfavorable Financial Terms and Conditions in Shopping Websites at Scale
- What Do Position Embeddings Learn? An Empirical Study of Pre-Trained Language Model Positional Encoding
- Scientific Claim Verification with VERT5ERINI
- Discovering Useful Sentence Representations from Large Pretrained Language Models
- Towards Socially Intelligent Agents with Mental State Transition and Human Utility
- DORA: Toward Policy Optimization for Task-oriented Dialogue System with Efficient Context
- ERNIE-Tiny : A Progressive Distillation Framework for Pretrained Transformer Compression
- Evaluating the Reliability of Self-Explanations in Large Language Models
- Semi-Supervised Class Discovery
- LARR: Large Language Model Aided Real-time Scene Recommendation with Semantic Understanding
- Generative Emotion Cause Explanation in Multimodal Conversations
- Advances in Multi-turn Dialogue Comprehension: A Survey
- Improving the Efficiency of Grammatical Error Correction with Erroneous Span Detection and Correction
- Creating a silver standard for patent simplification
- SuperShaper: Task-Agnostic Super Pre-training of BERT Models with Variable Hidden Dimensions
- Studying the role of named entities for content preservation in text style transfer
- Neural Language Modeling for Contextualized Temporal Graph Generation
- COBE: Contextualized Object Embeddings from Narrated Instructional Video
- El Departamento de Nosotros: How Machine Translated Corpora Affects Language Models in MRC Tasks
- EL-Attention: Memory Efficient Lossless Attention for Generation
- Recommending Variable Names for Extract Local Variable Refactorings
- Robustly Optimized and Distilled Training for Natural Language Understanding
- Conversational Answer Generation and Factuality for Reading Comprehension Question-Answering
- Transfer Learning of Transformer-based Speech Recognition Models from Czech to Slovak
- Directed Beam Search: Plug-and-Play Lexically Constrained Language Generation
- Generative Meta-Learning for Zero-Shot Relation Triplet Extraction
- Improving Task-Agnostic BERT Distillation with Layer Mapping Search
- Current Limitations of Language Models: What You Need is Retrieval
- M2DS: Multilingual Dataset for Multi-document Summarisation
- Comparison of Czech Transformers on Text Classification Tasks
- Semantic Search as Extractive Paraphrase Span Detection
- MemoNet: Memorizing All Cross Features' Representations Efficiently via Multi-Hash Codebook Network for CTR Prediction
- Continual Mixed-Language Pre-Training for Extremely Low-Resource Neural Machine Translation
- An Overview on Generative AI at Scale with Edge-Cloud Computing
- A First Look: Towards Explainable TextVQA Models via Visual and Textual Explanations
- How Important is Domain Specificity in Language Models and Instruction Finetuning for Biomedical Relation Extraction?
- I've Got 99 Problems But FLOPS Ain't One
- Real-Time Human Action Recognition on Embedded Platforms
- Embedding an Ethical Mind: Aligning Text-to-Image Synthesis via Lightweight Value Optimization
- Wrong Answers Can Also Be Useful: PlausibleQA -- A Large-Scale QA Dataset with Answer Plausibility Scores
- Valid Text-to-SQL Generation with Unification-based DeepStochLog
- Enhancing Ultra-Low-Bit Quantization of Large Language Models Through Saliency-Aware Partial Retraining
- Hierarchical Spatial-Temporal Graph-Enhanced Model for Map-Matching
- PLM4NDV: Minimizing Data Access for Number of Distinct Values Estimation with Pre-trained Language Models
- Iterative Self-Training for Code Generation via Reinforced Re-Ranking
- DISCIE -- Discriminative Closed Information Extraction
- Analyzing the Influence of Knowledge Graph Information on Relation Extraction
- Boosting Explainability through Selective Rationalization in Pre-trained Language Models
- Zero-Shot vs. Few-Shot Multi-Speaker TTS Using Pre-trained Czech SpeechT5 Model
- Improving the quality of Persian clinical text with a novel spelling correction system
- Local Knowledge Powered Conversational Agents
- MSG-Chart: Multimodal Scene Graph for ChartQA
- Rating Prediction in Conversational Task Assistants with Behavioral and Conversational-Flow Features
- Adversarial Self-Supervised Data-Free Distillation for Text Classification
- CCAE: A Corpus of Chinese-based Asian Englishes
- In-Context Learning for Knowledge Base Question Answering for Unmanned Systems based on Large Language Models
- Contrastive Fine-tuning Improves Robustness for Neural Rankers
- Narrative Question Answering with Cutting-Edge Open-Domain QA Techniques: A Comprehensive Study
- Super Tickets in Pre-Trained Language Models: From Model Compression to Improving Generalization
- Interpretable Self-supervised Multi-task Learning for COVID-19 Information Retrieval and Extraction
- A Simple and Interpretable Predictive Model for Healthcare
- Automatic Claim Review for Climate Science via Explanation Generation
- MetaXT: Meta Cross-Task Transfer between Disparate Label Spaces
- Leveraging Pretrained Models for Automatic Summarization of Doctor-Patient Conversations
- Composed Fine-Tuning: Freezing Pre-Trained Denoising Autoencoders for Improved Generalization
- Multiplicative Position-aware Transformer Models for Language Understanding
- Dense Hierarchical Retrieval for Open-Domain Question Answering
- Gestalt: a Stacking Ensemble for SQuAD2.0
- BERT-JAM: Boosting BERT-Enhanced Neural Machine Translation with Joint Attention
- WaLDORf: Wasteless Language-model Distillation On Reading-comprehension
- Digging Deeper into CRNN Model in Chinese Text Images Recognition
- Learning to Emphasize: Dataset and Shared Task Models for Selecting Emphasis in Presentation Slides
- A Road Map to Strong Intelligence
- Self-supervised Text-to-SQL Learning with Header Alignment Training
- A Review on Semi-Supervised Relation Extraction
- Self-supervised Regularization for Text Classification
- Probabilistically Masked Language Model Capable of Autoregressive Generation in Arbitrary Word Order
- General Purpose Text Embeddings from Pre-trained Language Models for Scalable Inference
- Towards QoS-Aware and Resource-Efficient GPU Microservices Based on Spatial Multitasking GPUs In Datacenters
- G5: A Universal GRAPH-BERT for Graph-to-Graph Transfer and Apocalypse Learning
- ANA at SemEval-2020 Task 4: mUlti-task learNIng for cOmmonsense reasoNing (UNION)
- Zero-Shot Open-Book Question Answering
- Cut the CARP: Fishing for zero-shot story evaluation
- Cheap Learning: Maximising Performance of Language Models for Social Data Science Using Minimal Data
- Structure-aware Fine-tuning of Sequence-to-sequence Transformers for Transition-based AMR Parsing
- Autoregressive Knowledge Distillation through Imitation Learning
- Incorporating Behavioral Hypotheses for Query Generation
- TransAug: Translate as Augmentation for Sentence Embeddings
- Fine-tuning Multi-hop Question Answering with Hierarchical Graph Network
- CHIME: Cross-passage Hierarchical Memory Network for Generative Review Question Answering
- Converting the Point of View of Messages Spoken to Virtual Assistants
- RUEL: Retrieval-Augmented User Representation with Edge Browser Logs for Sequential Recommendation
- Comparison of pipeline, sequence-to-sequence, and GPT models for end-to-end relation extraction: experiments with the rare disease use-case
- Contextualization with SPLADE for High Recall Retrieval
- Can questions summarize a corpus? Using question generation for characterizing COVID-19 research
- SHARE: a System for Hierarchical Assistive Recipe Editing
- Searching Personal Collections
- Spline-based Transformers
- Normalization of Input-output Shared Embeddings in Text Generation Models
- Normalizador Neural de Datas e Endereços
- Analyzing Curriculum Learning for Sentiment Analysis along Task Difficulty, Pacing and Visualization Axes
- T-EMDE: Sketching-based global similarity for cross-modal retrieval
- GRPO and Reflection Reward for Mathematical Reasoning in Large Language Models
- Benchmarking down-scaled (not so large) pre-trained language models
- Incorporating Commonsense Knowledge Graph in Pretrained Models for Social Commonsense Tasks
- Automatic Generation of Highlights for Academic Paper Via Prompt-based Learning
- Iterative Decoding for Compositional Generalization in Transformers
- One Network Fits All? Modular versus Monolithic Task Formulations in Neural Networks
- Why Can You Lay Off Heads? Investigating How BERT Heads Transfer
- Matryoshka Model Learning for Improved Elastic Student Models
- DESCGEN: A Distantly Supervised Dataset for Generating Abstractive Entity Descriptions
- Multi-Pair Text Style Transfer on Unbalanced Data
- Generating abstractive summaries of Lithuanian news articles using a transformer model
- Neural Supervised Domain Adaptation by Augmenting Pre-trained Models with Random Units
- Improve Query Focused Abstractive Summarization by Incorporating Answer Relevance
- DeepMutants: Training neural bug detectors with contextual mutations
- Inference Time Style Control for Summarization
- Learning Algebraic Recombination for Compositional Generalization
- Beyond In-Place Corruption: Insertion and Deletion In Denoising Probabilistic Models
- Efficient Retrieval Optimized Multi-task Learning
- QMUL-SDS at SCIVER: Step-by-Step Binary Classification for Scientific Claim Verification
- TIAGE: A Benchmark for Topic-Shift Aware Dialog Modeling
- Exploring a Unified Sequence-To-Sequence Transformer for Medical Product Safety Monitoring in Social Media
- Capturing Structural Locality in Non-parametric Language Models
- How BPE Affects Memorization in Transformers
- The Efficiency Misnomer
- Answer Generation for Questions With Multiple Information Sources in E-Commerce
- STIL -- Simultaneous Slot Filling, Translation, Intent Classification, and Language Identification: Initial Results using mBART on MultiATIS++
- Large Product Key Memory for Pretrained Language Models
- PARADE: A New Dataset for Paraphrase Identification Requiring Computer Science Domain Knowledge
- When in Doubt, Summon the Titans: Efficient Inference with Large Models