Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks
arXiv:1506.03099
Abstract
Recurrent Neural Networks can be trained to produce sequences of tokens given some input, as exemplified by recent results in machine translation and image captioning. The current approach to training them consists of maximizing the likelihood of each token in the sequence given the current (recurrent) state and the previous token. At inference, the unknown previous token is then replaced by a token generated by the model itself. This discrepancy between training and inference can yield errors that can accumulate quickly along the generated sequence. We propose a curriculum learning strategy to gently change the training process from a fully guided scheme using the true previous token, towards a less guided scheme which mostly uses the generated token instead. Experiments on several sequence prediction tasks show that this approach yields significant improvements. Moreover, it was used successfully in our winning entry to the MSCOCO image captioning challenge, 2015.
References in corpus (5)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Sequence to Sequence Learning with Neural Networks
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- End-to-end Continuous Speech Recognition using Attention-based Recurrent NN: First Results
- Grammar as a Foreign Language
Cited by in corpus (420)
- Survey of Hallucination in Natural Language Generation
- Recent Trends in Deep Learning Based Natural Language Processing
- Sequence Level Training with Recurrent Neural Networks
- Show and Tell: Lessons learned from the 2015 MSCOCO Image Captioning Challenge
- FastSpeech: Fast, Robust and Controllable Text to Speech
- Diffusion Convolutional Recurrent Neural Network: Data-Driven Traffic Forecasting
- Cascaded Diffusion Models for High Fidelity Image Generation
- Improved Image Captioning via Policy Gradient optimization of SPIDEr
- Deep Predictive Coding Networks for Video Prediction and Unsupervised Learning
- Model-Based Reinforcement Learning for Atari
- BRITS: Bidirectional Recurrent Imputation for Time Series
- LSTM-based Encoder-Decoder for Multi-sensor Anomaly Detection
- Professor Forcing: A New Algorithm for Training Recurrent Networks
- Listen, Attend and Spell
- A Hierarchical Latent Vector Model for Learning Long-Term Structure in Music
- SeqGAN: Sequence Generative Adversarial Nets with Policy Gradient
- Accelerating Eulerian Fluid Simulation With Convolutional Networks
- How to Build a Graph-Based Deep Learning Architecture in Traffic Domain: A Survey
- Character Controllers Using Motion VAEs
- Unsupervised Learning for Physical Interaction through Video Prediction
- Traffic Prediction using Artificial Intelligence: Review of Recent Advances and Emerging Opportunities
- A Multi-Horizon Quantile Recurrent Forecaster
- Natural Language Processing Advancements By Deep Learning: A Survey
- How (not) to Train your Generative Model: Scheduled Sampling, Likelihood, Adversary?
- End-to-End Relation Extraction using LSTMs on Sequences and Tree Structures
- QuaterNet: A Quaternion-based Recurrent Model for Human Motion
- A Survey on Neural Speech Synthesis
- ST-GRAT: A Novel Spatio-temporal Graph Attention Network for Accurately Forecasting Dynamically Changing Road Speed
- Dynamical Variational Autoencoders: A Comprehensive Review
- Multi-Sensor Prognostics using an Unsupervised Health Index based on LSTM Encoder-Decoder
- Long Text Generation via Adversarial Training with Leaked Information
- MOReL : Model-Based Offline Reinforcement Learning
- Latent ODEs for Irregularly-Sampled Time Series
- Survey on reinforcement learning for language processing
- Modeling Human Motion with Quaternion-based Neural Networks
- Tacotron: Towards End-to-End Speech Synthesis
- Automatic Goal Generation for Reinforcement Learning Agents
- GraphVAE: Towards Generation of Small Graphs Using Variational Autoencoders
- A Convolutional Attention Network for Extreme Summarization of Source Code
- Reverse Curriculum Generation for Reinforcement Learning
- Neural Summarization by Extracting Sentences and Words
- Neural Machine Translation and Sequence-to-sequence Models: A Tutorial
- Land Cover Classification from Multi-temporal, Multi-spectral Remotely Sensed Imagery using Patch-Based Recurrent Neural Networks
- Recurrent Environment Simulators
- Zero-Shot Out-of-Distribution Detection Based on the Pre-trained Model CLIP
- Sequence-to-Sequence Learning as Beam-Search Optimization
- A survey on GANs for computer vision: Recent research, analysis and taxonomy
- Generating News Headlines with Recurrent Neural Networks
- Recurrent Dropout without Memory Loss
- Listen and Fill in the Missing Letters: Non-Autoregressive Transformer for Speech Recognition
- Reward Augmented Maximum Likelihood for Neural Structured Prediction
- Video Paragraph Captioning Using Hierarchical Recurrent Neural Networks
- Deeply AggreVaTeD: Differentiable Imitation Learning for Sequential Prediction
- Convolutional Tensor-Train LSTM for Spatio-temporal Learning
- Language GANs Falling Short
- Machine Learning for Spatiotemporal Sequence Forecasting: A Survey
- On Accurate Evaluation of GANs for Language Generation
- RETURNN as a Generic Flexible Neural Toolkit with Application to Translation and Speech Recognition
- Improving Image Captioning with Conditional Generative Adversarial Nets
- Language-guided Navigation via Cross-Modal Grounding and Alternate Adversarial Learning
- A Graph to Graphs Framework for Retrosynthesis Prediction
- Long-term Forecasting using Higher Order Tensor RNNs
- NetTraj: A Network-based Vehicle Trajectory Prediction Model with Directional Representation and Spatiotemporal Attention Mechanisms
- Decoupled Novel Object Captioner
- Graph Neural Networks for Natural Language Processing: A Survey
- Retrieval-Augmented Generation for Code Summarization via Hybrid GNN
- Neural Abstractive Text Summarization with Sequence-to-Sequence Models
- Conditional GAN for timeseries generation
- Interpretable Sequence Learning for COVID-19 Forecasting
- A Game Theoretic Framework for Model Based Reinforcement Learning
- DP-GAN: Diversity-Promoting Generative Adversarial Network for Generating Informative and Diversified Text
- Coordinated Multi-Agent Imitation Learning
- Self-supervised Learning for Video Correspondence Flow
- Neural Text Generation: Past, Present and Beyond
- Dance Revolution: Long-Term Dance Generation with Music via Curriculum Learning
- ImageBART: Bidirectional Context with Multinomial Diffusion for Autoregressive Image Synthesis
- Forecast Network-Wide Traffic States for Multiple Steps Ahead: A Deep Learning Approach Considering Dynamic Non-Local Spatial Correlation and Non-Stationary Temporal Dependency
- Attention on Attention for Image Captioning
- A Unified Query-based Generative Model for Question Generation and Question Answering
- Noisy Parallel Approximate Decoding for Conditional Recurrent Language Model
- Reinforcement Learning Based Graph-to-Sequence Model for Natural Question Generation
- Multiway Non-rigid Point Cloud Registration via Learned Functional Map Synchronization
- Seq2Slate: Re-ranking and Slate Optimization with RNNs
- Learning Global Features for Coreference Resolution
- End-to-End Dense Video Captioning with Masked Transformer
- Calibration of Encoder Decoder Models for Neural Machine Translation
- VATEX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research
- From Language to Programs: Bridging Reinforcement Learning and Maximum Marginal Likelihood
- Look, Imagine and Match: Improving Textual-Visual Cross-Modal Retrieval with Generative Models
- A Neural Transducer
- LipSound2: Self-Supervised Pre-Training for Lip-to-Speech Reconstruction and Lip Reading
- Time Series Data Imputation: A Survey on Deep Learning Approaches
- End-to-End Adversarial Text-to-Speech
- A Fast Unified Model for Parsing and Sentence Understanding
- Motion Generation Using Bilateral Control-Based Imitation Learning with Autoregressive Learning
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++
- Feature-aware conditional GAN for category text generation
- A Tutorial on Deep Latent Variable Models of Natural Language
- Improving Variational Encoder-Decoders in Dialogue Generation
- Speaking the Same Language: Matching Machine to Human Captions by Adversarial Training
- Show, Adapt and Tell: Adversarial Training of Cross-domain Image Captioner
- A Paragraph-level Multi-task Learning Model for Scientific Fact-Verification
- VideoFlow: A Conditional Flow-Based Model for Stochastic Video Generation
- FitVid: Overfitting in Pixel-Level Video Prediction
- Bridging the Gap between Training and Inference for Neural Machine Translation
- Non-autoregressive Transformer-based End-to-end ASR using BERT
- PECOS: Prediction for Enormous and Correlated Output Spaces
- Incorporating Global Visual Features into Attention-Based Neural Machine Translation
- Combating the Compounding-Error Problem with a Multi-step Model
- Toward Diverse Text Generation with Inverse Reinforcement Learning
- Regularizing RNNs for Caption Generation by Reconstructing The Past with The Present
- Task Loss Estimation for Sequence Prediction
- MPC-Inspired Neural Network Policies for Sequential Decision Making
- Less Is More: Picking Informative Frames for Video Captioning
- Masked Non-Autoregressive Image Captioning
- Learning to Segment Inputs for NMT Favors Character-Level Processing
- Model-based Reinforcement Learning for Semi-Markov Decision Processes with Neural ODEs
- Multilingual Language Processing From Bytes
- Improving Sequence-to-Sequence Learning via Optimal Transport
- Music SketchNet: Controllable Music Generation via Factorized Representations of Pitch and Rhythm
- Informative Visual Storytelling with Cross-modal Rules
- Neural Language Generation: Formulation, Methods, and Evaluation
- Improving Generalization of Transformer for Speech Recognition with Parallel Schedule Sampling and Relative Positional Embedding
- Video Captioning via Hierarchical Reinforcement Learning
- Memory In Memory: A Predictive Neural Network for Learning Higher-Order Non-Stationarity from Spatiotemporal Dynamics
- Cold-Start Reinforcement Learning with Softmax Policy Gradient
- Parallel Tacotron: Non-Autoregressive and Controllable TTS
- Paraphrase Generation with Deep Reinforcement Learning
- ScreenerNet: Learning Self-Paced Curriculum for Deep Neural Networks
- Hierarchically Structured Reinforcement Learning for Topically Coherent Visual Story Generation
- YouTube-VOS: Sequence-to-Sequence Video Object Segmentation
- Training with Exploration Improves a Greedy Stack-LSTM Parser
- Distilling Knowledge Learned in BERT for Text Generation
- SemStyle: Learning to Generate Stylised Image Captions using Unaligned Text
- Online Automatic Speech Recognition with Listen, Attend and Spell Model
- Prediction and Control with Temporal Segment Models
- MoGERNN: An Inductive Traffic Predictor for Unobserved Locations
- Next-Step Conditioned Deep Convolutional Neural Networks Improve Protein Secondary Structure Prediction
- A machine learning and feature engineering approach for the prediction of the uncontrolled re-entry of space objects
- Multi-Label Learning from Medical Plain Text with Convolutional Residual Models
- Unsupervised Learning of Object Structure and Dynamics from Videos
- Learning Dynamics Model in Reinforcement Learning by Incorporating the Long Term Future
- Fast Transient Simulation of High-Speed Channels Using Recurrent Neural Network
- Generating Emotive Gaits for Virtual Agents Using Affect-Based Autoregression
- Differentiable Top-k Operator with Optimal Transport
- Multi-Range Attentive Bicomponent Graph Convolutional Network for Traffic Forecasting
- Operator Learning with Neural Fields: Tackling PDEs on General Geometries
- SALSA-TEXT : self attentive latent space based adversarial text generation
- Neural Guided Constraint Logic Programming for Program Synthesis
- Watch, Listen, and Describe: Globally and Locally Aligned Cross-Modal Attentions for Video Captioning
- From Audio to Semantics: Approaches to end-to-end spoken language understanding
- Evolving Graphical Planner: Contextual Global Planning for Vision-and-Language Navigation
- On Exposure Bias, Hallucination and Domain Shift in Neural Machine Translation
- Espresso: A Fast End-to-end Neural Speech Recognition Toolkit
- Continuous PDE Dynamics Forecasting with Implicit Neural Representations
- ATOM: Commit Message Generation Based on Abstract Syntax Tree and Hybrid Ranking
- Robust Navigation with Language Pretraining and Stochastic Sampling
- Joint Modeling of Local and Global Temporal Dynamics for Multivariate Time Series Forecasting with Missing Values
- An Empirical Comparison on Imitation Learning and Reinforcement Learning for Paraphrase Generation
- SEARNN: Training RNNs with Global-Local Losses
- Learning audio sequence representations for acoustic event classification
- Eliminating Exposure Bias and Loss-Evaluation Mismatch in Multiple Object Tracking
- Natural Question Generation with Reinforcement Learning Based Graph-to-Sequence Model
- Neural-Guided Symbolic Regression with Asymptotic Constraints
- Connecting the Dots Between MLE and RL for Sequence Prediction
- Clockwork Variational Autoencoders
- Joint Air Quality and Weather Prediction Based on Multi-Adversarial Spatiotemporal Networks
- Guided Dialog Policy Learning without Adversarial Learning in the Loop
- Controlling Computation versus Quality for Neural Sequence Models
- Parallel Scheduled Sampling
- A Comprehensive Survey of Deep Learning for Image Captioning
- Temporal Difference Variational Auto-Encoder
- Multimodal Machine Translation with Reinforcement Learning
- Jointly Measuring Diversity and Quality in Text Generation Models
- Sentiment-oriented Transformer-based Variational Autoencoder Network for Live Video Commenting
- Gated Embeddings in End-to-End Speech Recognition for Conversational-Context Fusion
- Self-Adversarial Learning with Comparative Discrimination for Text Generation
- Dynamic Hybrid Relation Network for Cross-Domain Context-Dependent Semantic Parsing
- Improving the Performance of Automated Audio Captioning via Integrating the Acoustic and Semantic Information
- Improved Training with Curriculum GANs
- Variable-rate discrete representation learning
- End-to-End Reinforcement Learning of Koopman Models for Economic Nonlinear Model Predictive Control
- A Brief Introduction to Generative Models
- Sensor-Augmented Egocentric-Video Captioning with Dynamic Modal Attention
- Discriminative Adversarial Search for Abstractive Summarization
- Contrastive Learning with Adversarial Perturbations for Conditional Text Generation
- A Survey on Curriculum Learning
- ALBA : Reinforcement Learning for Video Object Segmentation
- Multivariate, Multistep Forecasting, Reconstruction and Feature Selection of Ocean Waves via Recurrent and Sequence-to-Sequence Networks
- Hard but Robust, Easy but Sensitive: How Encoder and Decoder Perform in Neural Machine Translation
- Attention based end to end Speech Recognition for Voice Search in Hindi and English
- Regularizing Neural Machine Translation by Target-bidirectional Agreement
- Structure-Aware Abstractive Conversation Summarization via Discourse and Action Graphs
- Polyphonic Piano Transcription Using Autoregressive Multi-State Note Model
- Scheduled Sampling in Vision-Language Pretraining with Decoupled Encoder-Decoder Network
- Topic-Preserving Synthetic News Generation: An Adversarial Deep Reinforcement Learning Approach
- BERT as a Teacher: Contextual Embeddings for Sequence-Level Reward
- Discourse-Aware Neural Rewards for Coherent Text Generation
- Interpretable Deep Feature Propagation for Early Action Recognition
- A Multi-task Learning Framework for Drone State Identification and Trajectory Prediction
- Improving Conditional Sequence Generative Adversarial Networks by Stepwise Evaluation
- Beyond Error Propagation in Neural Machine Translation: Characteristics of Language Also Matter
- Image Captioning Based on a Hierarchical Attention Mechanism and Policy Gradient Optimization
- ReWE: Regressing Word Embeddings for Regularization of Neural Machine Translation Systems
- Class LM and word mapping for contextual biasing in End-to-End ASR
- Neural Machine Translation with Adequacy-Oriented Learning
- Deep Model-Based Reinforcement Learning for High-Dimensional Problems, a Survey
- An Features Extraction and Recognition Method for Underwater Acoustic Target Based on ATCNN
- Learning representations for multivariate time series with missing data using Temporal Kernelized Autoencoders
- What Makes A Good Story? Designing Composite Rewards for Visual Storytelling
- Rethinking Exposure Bias In Language Modeling
- Multi-modal Feature Fusion with Feature Attention for VATEX Captioning Challenge 2020
- TwoStreamVAN: Improving Motion Modeling in Video Generation
- AutoAssist: A Framework to Accelerate Training of Deep Neural Networks
- Attention-based Transducer for Online Speech Recognition
- Decision-Aware Conditional GANs for Time Series Data
- Bilingual-GAN: A Step Towards Parallel Text Generation
- Greedy Search with Probabilistic N-gram Matching for Neural Machine Translation
- Platoon trajectories generation: A unidirectional interconnected LSTM-based car following model
- CIF: Continuous Integrate-and-Fire for End-to-End Speech Recognition
- A Two-part Transformer Network for Controllable Motion Synthesis
- Utilizing Character and Word Embeddings for Text Normalization with Sequence-to-Sequence Models
- Adaptive Skip Intervals: Temporal Abstraction for Recurrent Dynamical Models
- Finding It at Another Side: A Viewpoint-Adapted Matching Encoder for Change Captioning
- Delving Deeper into the Decoder for Video Captioning
- End-to-End Video Captioning with Multitask Reinforcement Learning
- Amharic Abstractive Text Summarization
- An Online Attention-based Model for Speech Recognition
- A New GAN-based End-to-End TTS Training Algorithm
- Uncertainty Prediction for Deep Sequential Regression Using Meta Models
- Audio-Linguistic Embeddings for Spoken Sentences
- ARAML: A Stable Adversarial Training Framework for Text Generation
- Training Stronger Baselines for Learning to Optimize
- Better Captioning with Sequence-Level Exploration
- Generating Music with a Self-Correcting Non-Chronological Autoregressive Model
- Visual Storytelling via Predicting Anchor Word Embeddings in the Stories
- Document-level Event Extraction via Heterogeneous Graph-based Interaction Model with a Tracker
- Unpaired Cross-lingual Image Caption Generation with Self-Supervised Rewards
- Judge the Judges: A Large-Scale Evaluation Study of Neural Language Models for Online Review Generation
- Cross-Utterance Language Models with Acoustic Error Sampling
- Optimal Completion Distillation for Sequence Learning
- Preventing Posterior Collapse with Levenshtein Variational Autoencoder
- CatVRNN: Generating Category Texts via Multi-task Learning
- Non-parallel Voice Conversion System with WaveNet Vocoder and Collapsed Speech Suppression
- Text Generation by Learning from Demonstrations
- Multilingual End-to-End Speech Translation
- Translating Navigation Instructions in Natural Language to a High-Level Plan for Behavioral Robot Navigation
- A comprehensive study of batch construction strategies for recurrent neural networks in MXNet
- Rescuing neural spike train models from bad MLE
- Who did They Respond to? Conversation Structure Modeling using Masked Hierarchical Transformer
- FDNet: A Deep Learning Approach with Two Parallel Cross Encoding Pathways for Precipitation Nowcasting
- Out-of-Distribution Dynamics Detection: RL-Relevant Benchmarks and Results
- Deep Probabilistic Time Series Forecasting using Augmented Recurrent Input for Dynamic Systems
- Detecting and Exorcising Statistical Demons from Language Models with Anti-Models of Negative Data
- Time Series Motion Generation Considering Long Short-Term Motion
- Non-Autoregressive Translation by Learning Target Categorical Codes
- Doc2EDAG: An End-to-End Document-level Framework for Chinese Financial Event Extraction
- Learning to Compose Topic-Aware Mixture of Experts for Zero-Shot Video Captioning
- Distilling the Knowledge of BERT for Sequence-to-Sequence ASR
- Investigating Linguistic Pattern Ordering in Hierarchical Natural Language Generation
- Closed-Book Training to Improve Summarization Encoder Memory
- Attention Forcing for Sequence-to-sequence Model Training
- Mollifying Networks
- Regularizing Dialogue Generation by Imitating Implicit Scenarios
- Distilling Knowledge for Search-based Structured Prediction
- Energy-Based Models with Applications to Speech and Language Processing
- Text Generation Based on Generative Adversarial Nets with Latent Variable
- Multi-Head Adapter Routing for Cross-Task Generalization
- Residual Energy-Based Models for End-to-End Speech Recognition
- Teacher-Student Training for Robust Tacotron-based TTS
- B-SCST: Bayesian Self-Critical Sequence Training for Image Captioning
- From Credit Assignment to Entropy Regularization: Two New Algorithms for Neural Sequence Prediction
- Recent Progresses in Deep Learning based Acoustic Models (Updated)
- Attention Guided Dialogue State Tracking with Sparse Supervision
- Spatial Memory for Context Reasoning in Object Detection
- Customizing Sequence Generation with Multi-Task Dynamical Systems
- MC-LSTM: Mass-Conserving LSTM
- End-to-End Monaural Multi-speaker ASR System without Pretraining
- Alleviating the Burden of Labeling: Sentence Generation by Attention Branch Encoder-Decoder Network
- Attend More Times for Image Captioning
- DSS: Synthesizing long Digital Ink using Data augmentation, Style encoding and Split generation
- Bayesian surrogate learning in dynamic simulator-based regression problems
- Bridging Subword Gaps in Pretrain-Finetune Paradigm for Natural Language Generation
- Adding A Filter Based on The Discriminator to Improve Unconditional Text Generation
- Action Search: Spotting Actions in Videos and Its Application to Temporal Action Localization
- Earlier Isn't Always Better: Sub-aspect Analysis on Corpus and System Biases in Summarization
- Convolutional Recurrent Reconstructive Network for Spatiotemporal Anomaly Detection in Solder Paste Inspection
- End-to-End Automatic Speech Recognition with Deep Mutual Learning
- Collaborative Training of GANs in Continuous and Discrete Spaces for Text Generation
- Learning Disentangled Representations of Video with Missing Data
- Controllable Dual Skew Divergence Loss for Neural Machine Translation
- Adversarial Sub-sequence for Text Generation
- Learning Word-Level Confidence For Subword End-to-End ASR
- Alleviate Exposure Bias in Sequence Prediction \\ with Recurrent Neural Networks
- Geometry-Entangled Visual Semantic Transformer for Image Captioning
- Fast DCTTS: Efficient Deep Convolutional Text-to-Speech
- Attention Forcing for Machine Translation
- Adversarial Neural Trip Recommendation
- Hierarchical Multi-Grained Generative Model for Expressive Speech Synthesis
- Improving Text Generation with Student-Forcing Optimal Transport
- Promising Accurate Prefix Boosting for sequence-to-sequence ASR
- Modular Action Concept Grounding in Semantic Video Prediction
- Learning to Explicitate Connectives with Seq2Seq Network for Implicit Discourse Relation Classification
- WSRNet: Joint Spotting and Recognition of Handwritten Words
- STYLER: Style Factor Modeling with Rapidity and Robustness via Speech Decomposition for Expressive and Controllable Neural Text to Speech
- Reinforced Generative Adversarial Network for Abstractive Text Summarization
- Connecting What to Say With Where to Look by Modeling Human Attention Traces
- An End-to-End Neural Network for Polyphonic Piano Music Transcription
- Chained Predictions Using Convolutional Neural Networks
- Language Model Evaluation in Open-ended Text Generation
- Modularized Textual Grounding for Counterfactual Resilience
- Melody Structure Transfer Network: Generating Music with Separable Self-Attention
- Neural Machine Translation: A Review and Survey
- Adaptive Correlated Monte Carlo for Contextual Categorical Sequence Generation
- Forward-Backward Decoding for Regularizing End-to-End TTS
- Writing Polishment with Simile: Task, Dataset and A Neural Approach
- Quantum Optical Experiments Modeled by Long Short-Term Memory
- Improving the Robustness to Data Inconsistency between Training and Testing for Code Completion by Hierarchical Language Model
- Boosted Attention: Leveraging Human Attention for Image Captioning
- Improve Language Modelling for Code Completion through Statement Level Language Model based on Statement Embedding Generated by BiLSTM
- Learning to Generate 3D Shapes with Generative Cellular Automata
- Deep Recurrent Encoder: A scalable end-to-end network to model brain signals
- Graph Neural Network based Service Function Chaining for Automatic Network Control
- Sequence-to-Sequence Learning on Keywords for Efficient FAQ Retrieval
- Dynamic Relational Inference in Multi-Agent Trajectories
- HUMBO: Bridging Response Generation and Facial Expression Synthesis
- Policy Gradients Incorporating the Future
- What Averages Do Not Tell -- Predicting Real Life Processes with Sequential Deep Learning
- Learning Shared Encoding Representation for End-to-End Speech Recognition Models
- Handwriting styles: benchmarks and evaluation metrics
- A Differentially Private Multi-Output Deep Generative Networks Approach For Activity Diary Synthesis
- Comparing Task Simplifications to Learn Closed-Loop Object Picking Using Deep Reinforcement Learning
- Audio-Visual Decision Fusion for WFST-based and seq2seq Models
- Making Classical Machine Learning Pipelines Differentiable: A Neural Translation Approach
- Deep Shallow Fusion for RNN-T Personalization
- Internal Language Model Training for Domain-Adaptive End-to-End Speech Recognition
- Differentiable Sampling with Flexible Reference Word Order for Neural Machine Translation
- Bridging the Gap Between Training and Inference for Spatio-Temporal Forecasting
- Mathematical Word Problem Generation from Commonsense Knowledge Graph and Equations
- Game of GANs: Game-Theoretical Models for Generative Adversarial Networks
- The MIDI Degradation Toolkit: Symbolic Music Augmentation and Correction
- Goal-driven text descriptions for images
- Distributional Discrepancy: A Metric for Unconditional Text Generation
- To Beam Or Not To Beam: That is a Question of Cooperation for Language GANs
- Weakly-Supervised Neural Response Selection from an Ensemble of Task-Specialised Dialogue Agents
- Phoneme Based Neural Transducer for Large Vocabulary Speech Recognition
- Deep Learning on Attributed Graphs: A Journey from Graphs to Their Embeddings and Back
- -Neighbor Based Curriculum Sampling for Sequence Prediction
- Curriculum Learning for Recurrent Video Object Segmentation
- Domain Adaptation via Teacher-Student Learning for End-to-End Speech Recognition
- Generalized Coarse-to-Fine Visual Recognition with Progressive Training
- Benchmarking Approximate Inference Methods for Neural Structured Prediction
- Learning to Stop in Structured Prediction for Neural Machine Translation
- Fixing exposure bias with imitation learning needs powerful oracles
- Protecting Anonymous Speech: A Generative Adversarial Network Methodology for Removing Stylistic Indicators in Text
- Image Captioning with Visual Object Representations Grounded in the Textual Modality
- Neural or Statistical: An Empirical Study on Language Models for Chinese Input Recommendation on Mobile
- State-Reification Networks: Improving Generalization by Modeling the Distribution of Hidden Representations
- Can We Automate Diagrammatic Reasoning?
- Character-Aware Attention-Based End-to-End Speech Recognition
- Can adversarial training learn image captioning ?
- VATEX Captioning Challenge 2019: Multi-modal Information Fusion and Multi-stage Training Strategy for Video Captioning
- Regressing Word and Sentence Embeddings for Regularization of Neural Machine Translation
- Simulation of an Elevator Group Control Using Generative Adversarial Networks and Related AI Tools
- Learn to Talk via Proactive Knowledge Transfer
- Conditioned Time-Dilated Convolutions for Sound Event Detection
- Dual Reinforcement-Based Specification Generation for Image De-Rendering
- Evaluation Discrepancy Discovery: A Sentence Compression Case-study
- Adversarial Machine Learning in Text Analysis and Generation
- Uncertainty-Aware Label Refinement for Sequence Labeling
- Autoregressive Reasoning over Chains of Facts with Transformers
- Reciprocal Supervised Learning Improves Neural Machine Translation
- Master Thesis: Neural Sign Language Translation by Learning Tokenization
- Autoregressive Knowledge Distillation through Imitation Learning
- Adaptive Bridge between Training and Inference for Dialogue
- Integrating Categorical Features in End-to-End ASR
- Mapping Language to Programs using Multiple Reward Components with Inverse Reinforcement Learning
- Nana-HDR: A Non-attentive Non-autoregressive Hybrid Model for TTS
- Learning Natural Language Generation from Scratch
- Transformer-based Lexically Constrained Headline Generation
- Factual Consistency Evaluation for Text Summarization via Counterfactual Estimation
- Scheduled Sampling Based on Decoding Steps for Neural Machine Translation
- Can the Transformer Be Used as a Drop-in Replacement for RNNs in Text-Generating GANs?
- Reducing Exposure Bias in Training Recurrent Neural Network Transducers
- Confidence-Aware Scheduled Sampling for Neural Machine Translation
- Surgical Instruction Generation with Transformers
- Single Model for Influenza Forecasting of Multiple Countries by Multi-task Learning
- Probabilistic Graph Reasoning for Natural Proof Generation
- A Conditional Splitting Framework for Efficient Constituency Parsing
- Recurrent Stacking of Layers in Neural Networks: An Application to Neural Machine Translation
- NAST: A Non-Autoregressive Generator with Word Alignment for Unsupervised Text Style Transfer
- Action2video: Generating Videos of Human 3D Actions
- Sequence-level self-learning with multiple hypotheses
- Mixture of Dynamical Variational Autoencoders for Multi-Source Trajectory Modeling and Separation
- FlowVOS: Weakly-Supervised Visual Warping for Detail-Preserving and Temporally Consistent Single-Shot Video Object Segmentation
- Towards Generating Real-World Time Series Data
- Symbolic Music Loop Generation with VQ-VAE
- Meta-Forecasting by combining Global Deep Representations with Local Adaptation
- TREND: Trigger-Enhanced Relation-Extraction Network for Dialogues
- Data-to-text Generation by Splicing Together Nearest Neighbors
- Self-attention aggregation network for video face representation and recognition
- Positioning yourself in the maze of Neural Text Generation: A Task-Agnostic Survey
- Spatial Attention as an Interface for Image Captioning Models
- ARMA Nets: Expanding Receptive Field for Dense Prediction
- Integrating Knowledge into End-to-End Speech Recognition from External Text-Only Data
- Improving OOV Detection and Resolution with External Language Models in Acoustic-to-Word ASR
- How Sequence-to-Sequence Models Perceive Language Styles?
- Exposure Bias versus Self-Recovery: Are Distortions Really Incremental for Autoregressive Text Generation?
- Taco-VC: A Single Speaker Tacotron based Voice Conversion with Limited Data
- PR Product: A Substitute for Inner Product in Neural Networks
- Error-Correcting Neural Sequence Prediction
- Generation of Synthetic Electronic Medical Record Text
- Dual Latent Variable Model for Low-Resource Natural Language Generation in Dialogue Systems
- Middle-Out Decoding
- Imitation Learning for Neural Morphological String Transduction
- Top-Down Tree Structured Text Generation
- Rethinking the Form of Latent States in Image Captioning
- Region Growing Curriculum Generation for Reinforcement Learning
- Automatic Documentation of ICD Codes with Far-Field Speech Recognition
- Learning the Reward Function for a Misspecified Model