Categorical Reparameterization with Gumbel-Softmax
arXiv:1611.01144
Abstract
Categorical variables are a natural choice for representing discrete structure in the world. However, stochastic neural networks rarely use categorical latent variables due to the inability to backpropagate through samples. In this work, we present an efficient gradient estimator that replaces the non-differentiable sample from a categorical distribution with a differentiable sample from a novel Gumbel-Softmax distribution. This distribution has the essential property that it can be smoothly annealed into a categorical distribution. We show that our Gumbel-Softmax estimator outperforms state-of-the-art gradient estimators on structured output prediction and unsupervised generative modeling tasks with categorical latent variables, and enables large speedups on semi-supervised classification.
Cited by in corpus (1031)
- An Introduction to Variational Autoencoders
- AutoML: A Survey of the State-of-the-Art
- Improved Training of Wasserstein GANs
- NIPS 2016 Tutorial: Generative Adversarial Networks
- Zero-Shot Text-to-Image Generation
- Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments
- BEiT: BERT Pre-Training of Image Transformers
- Contrastive Representation Learning: A Framework and Review
- Deep learning for molecular design - a review of the state of the art
- Unifying Knowledge Graph Learning and Recommendation: Towards a Better Understanding of User Preferences
- Style Transfer from Non-Parallel Text by Cross-Alignment
- MolGAN: An implicit generative model for small molecular graphs
- Multi-Task Learning with Deep Neural Networks: A Survey
- NVAE: A Deep Hierarchical Variational Autoencoder
- Generating Multi-label Discrete Patient Records using Generative Adversarial Networks
- Recent Advances in Autoencoder-Based Representation Learning
- DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification
- vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
- Pre-training Enhanced Spatial-temporal Graph Neural Network for Multivariate Time Series Forecasting
- Actor-Attention-Critic for Multi-Agent Reinforcement Learning
- A Review on Generative Adversarial Networks: Algorithms, Theory, and Applications
- Neural Relational Inference for Interacting Systems
- TensorFlow Distributions
- Online and Linear-Time Attention by Enforcing Monotonic Alignments
- Multi-Modal Self-Supervised Learning for Recommendation
- GANS for Sequences of Discrete Elements with the Gumbel-softmax Distribution
- Learning to Compose and Reason with Language Tree Structures for Visual Grounding
- Mixed Precision Quantization of ConvNets via Differentiable Neural Architecture Search
- NetGAN: Generating Graphs via Random Walks
- Stochastic Neural Networks for Hierarchical Reinforcement Learning
- Learning Discrete Structures for Graph Neural Networks
- Molecular graph generation with Graph Neural Networks
- Dynamic Convolutions: Exploiting Spatial Sparsity for Faster Inference
- Fast Decoding in Sequence Models using Discrete Latent Variables
- Bringing AI To Edge: From Deep Learning's Perspective
- Causal Confusion in Imitation Learning
- Learning Sparse Neural Networks through Regularization
- Adversarial Graph Augmentation to Improve Graph Contrastive Learning
- Plug and Play Language Models: A Simple Approach to Controlled Text Generation
- AdaShare: Learning What To Share For Efficient Deep Multi-Task Learning
- Network Pruning via Transformable Architecture Search
- GraphVAE: Towards Generation of Small Graphs Using Variational Autoencoders
- A survey on text generation using generative adversarial networks
- Sequential Latent Knowledge Selection for Knowledge-Grounded Dialogue
- Path-Restore: Learning Network Path Selection for Image Restoration
- A Disentangled Recognition and Nonlinear Dynamics Model for Unsupervised Learning
- Meta Learning Shared Hierarchies
- Gated-SCNN: Gated Shape CNNs for Semantic Segmentation
- Lifelong Generative Modeling
- Boundary-Seeking Generative Adversarial Networks
- FACMAC: Factored Multi-Agent Centralised Policy Gradients
- FedCP: Separating Feature Information for Personalized Federated Learning via Conditional Policy
- Learning to Generate Questions by Learning What not to Generate
- Learning to Explain: An Information-Theoretic Perspective on Model Interpretation
- Deep Neural Decision Trees
- Scalable Population Synthesis with Deep Generative Modeling
- Learning Disentangled Representations in the Imaging Domain
- MeanSum: A Neural Model for Unsupervised Multi-document Abstractive Summarization
- Generative Models for Automatic Chemical Design
- Graph Information Bottleneck
- An Interpretable Reasoning Network for Multi-Relation Question Answering
- Modeling Tabular data using Conditional GAN
- Improved Variational Autoencoders for Text Modeling using Dilated Convolutions
- Improving Fairness in Graph Neural Networks via Mitigating Sensitive Attribute Leakage
- Black-Box Attacks against RNN based Malware Detection Algorithms
- Backpropagation through the Void: Optimizing control variates for black-box gradient estimation
- On Explainability of Graph Neural Networks via Subgraph Explorations
- FBNet: Hardware-Aware Efficient ConvNet Design via Differentiable Neural Architecture Search
- Deep Structural Causal Models for Tractable Counterfactual Inference
- RNNLogic: Learning Logic Rules for Reasoning on Knowledge Graphs
- Disentangling Hate in Online Memes
- Best of Both Worlds: Transferring Knowledge from Discriminative Learning to a Generative Visual Dialog Model
- Low-Resource Knowledge-Grounded Dialogue Generation
- SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and Decomposition
- Learning Plannable Representations with Causal InfoGAN
- Phasebook and Friends: Leveraging Discrete Representations for Source Separation
- Jointly Learning Sentence Embeddings and Syntax with Unsupervised Tree-LSTMs
- NAS-Bench-1Shot1: Benchmarking and Dissecting One-shot Neural Architecture Search
- Adversarial Example Defenses: Ensembles of Weak Defenses are not Strong
- Sparse Sinkhorn Attention
- On Accurate Evaluation of GANs for Language Generation
- Concrete Autoencoders for Differentiable Feature Selection and Reconstruction
- Optimizing Molecules using Efficient Queries from Property Evaluations
- Non-Orthogonal Multiple Access Enhanced Multi-User Semantic Communication
- A review of Generative Adversarial Networks for Electronic Health Records: applications, evaluation measures and data sources
- Reducing Overestimation Bias in Multi-Agent Domains Using Double Centralized Critics
- In-situ crack and keyhole pore detection in laser directed energy deposition through acoustic signal and deep learning
- Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned
- Deep Unsupervised Learning for Joint Antenna Selection and Hybrid Beamforming
- Automated Machine Learning on Graphs: A Survey
- Discriminability objective for training descriptive captions
- Visually Grounded Neural Syntax Acquisition
- Relation Embedding with Dihedral Group in Knowledge Graph
- Anti-efficient encoding in emergent communication
- Diverse and Accurate Image Description Using a Variational Auto-Encoder with an Additive Gaussian Encoding Space
- Neural Speed Reading via Skim-RNN
- Modeling Others using Oneself in Multi-Agent Reinforcement Learning
- Fast differentiable DNA and protein sequence optimization for molecular design
- Temperature-transferable coarse-graining of ionic liquids with dual graph convolutional neural networks
- Lifelong Teacher-Student Network Learning
- A Probabilistic Formulation of Unsupervised Text Style Transfer
- Discovering Neural Wirings
- Learning Disentangled Joint Continuous and Discrete Representations
- Neural Oblivious Decision Ensembles for Deep Learning on Tabular Data
- Resource-efficient Deep Neural Networks for Automotive Radar Interference Mitigation
- DeCoAR 2.0: Deep Contextualized Acoustic Representations with Vector Quantization
- A Comprehensive Survey on Hardware-Aware Neural Architecture Search
- Improving Inference for Neural Image Compression
- Emergence of Language with Multi-agent Games: Learning to Communicate with Sequences of Symbols
- Theory and Experiments on Vector Quantized Autoencoders
- Generating Multi-Agent Trajectories using Programmatic Weak Supervision
- Variational inference with a quantum computer
- SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training
- Generative Discovery of Novel Chemical Designs using Diffusion Modeling and Transformer Deep Neural Networks with Application to Deep Eutectic Solvents
- Zero-Resource Knowledge-Grounded Dialogue Generation
- Emergent Translation in Multi-Agent Communication
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point Modeling
- A Wrapped Normal Distribution on Hyperbolic Space for Gradient-Based Learning
- Learning Generalisable Omni-Scale Representations for Person Re-Identification
- The continuous Bernoulli: fixing a pervasive error in variational autoencoders
- Mixed Precision DNNs: All you need is a good parametrization
- Winning the Lottery with Continuous Sparsification
- Learning Implicit Credit Assignment for Cooperative Multi-Agent Reinforcement Learning
- Semi-Supervised Variational Reasoning for Medical Dialogue Generation
- Actor-Critic based Training Framework for Abstractive Summarization
- Unsupervised Text Style Transfer using Language Models as Discriminators
- Explaining Question Answering Models through Text Generation
- Causal Discovery in Physical Systems from Videos
- Data Valuation using Reinforcement Learning
- Texture Memory-Augmented Deep Patch-Based Image Inpainting
- Learning with Differentiable Perturbed Optimizers
- SpotTune: Transfer Learning through Adaptive Fine-tuning
- Probabilistic Binary Neural Networks
- Generating Multi-Categorical Samples with Generative Adversarial Networks
- Optimizing Millions of Hyperparameters by Implicit Differentiation
- Parallel Attention Network with Sequence Matching for Video Grounding
- Off-Policy Multi-Agent Decomposed Policy Gradients
- Learning to Branch for Multi-Task Learning
- An Attention Free Transformer
- Creativity and Machine Learning: A Survey
- Robust Training of Vector Quantized Bottleneck Models
- Discrete Flows: Invertible Generative Models of Discrete Data
- Towards Binary-Valued Gates for Robust LSTM Training
- UCSG-Net -- Unsupervised Discovering of Constructive Solid Geometry Tree
- DADA: Differentiable Automatic Data Augmentation
- Dynamic Neural Networks: A Survey
- Learning to Evolve Structural Ensembles of Unfolded and Disordered Proteins Using Experimental Solution Data
- CASE: Learning Conditional Adversarial Skill Embeddings for Physics-based Characters
- UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled Data
- Multi-Agent Reinforcement Learning Based on Representational Communication for Large-Scale Traffic Signal Control
- Routing Networks and the Challenges of Modular and Compositional Computation
- Learning Permutations with Sinkhorn Policy Gradient
- DVAE++: Discrete Variational Autoencoders with Overlapping Transformations
- Adversarially Regularized Autoencoders
- Interpreting Graph Neural Networks for NLP With Differentiable Edge Masking
- Online Learned Continual Compression with Adaptive Quantization Modules
- Stochastic Optimization of Sorting Networks via Continuous Relaxations
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning
- Unifying Multimodal Transformer for Bi-directional Image and Text Generation
- Generative Adversarial Network Training is a Continual Learning Problem
- IDF++: Analyzing and Improving Integer Discrete Flows for Lossless Compression
- Inverse Graphics GAN: Learning to Generate 3D Shapes from Unstructured 2D Data
- ISTA-NAS: Efficient and Consistent Neural Architecture Search by Sparse Coding
- DisCoRL: Continual Reinforcement Learning via Policy Distillation
- Pixel-wise Attentional Gating for Parsimonious Pixel Labeling
- A Noise-Robust Self-supervised Pre-training Model Based Speech Representation Learning for Automatic Speech Recognition
- Knowledge-refined Denoising Network for Robust Recommendation
- Dropout Feature Ranking for Deep Learning Models
- LOREN: Logic-Regularized Reasoning for Interpretable Fact Verification
- Design by adaptive sampling
- The Functional Neural Process
- Emergent Multi-Agent Communication in the Deep Learning Era
- Unsupervised End-to-End Learning of Discrete Linguistic Units for Voice Conversion
- A Tutorial on Deep Latent Variable Models of Natural Language
- Stabilizing Invertible Neural Networks Using Mixture Models
- Airline Passenger Name Record Generation using Generative Adversarial Networks
- The Information Autoencoding Family: A Lagrangian Perspective on Latent Variable Generative Models
- Learning to Assemble Neural Module Tree Networks for Visual Grounding
- Fair DARTS: Eliminating Unfair Advantages in Differentiable Architecture Search
- Predicting Opinion Dynamics via Sociologically-Informed Neural Networks
- GraphDF: A Discrete Flow Model for Molecular Graph Generation
- Speaking the Same Language: Matching Machine to Human Captions by Adversarial Training
- On the interaction between supervision and self-play in emergent communication
- Improving the Gating Mechanism of Recurrent Neural Networks
- Greedy Attack and Gumbel Attack: Generating Adversarial Examples for Discrete Data
- Neural Nearest Neighbors Networks
- FBNetV2: Differentiable Neural Architecture Search for Spatial and Channel Dimensions
- Model-Based Counterfactual Synthesizer for Interpretation
- Multi-Task Deep Residual Echo Suppression with Echo-aware Loss
- Discrete and continuous representations and processing in deep learning: Looking forward
- A survey on Adversarial Recommender Systems: from Attack/Defense strategies to Generative Adversarial Networks
- Amortized Variational Inference: A Systematic Review
- Text as Neural Operator: Image Manipulation by Text Instruction
- Learning Multimodal Transition Dynamics for Model-Based Reinforcement Learning
- Learning to Denoise Biomedical Knowledge Graph for Robust Molecular Interaction Prediction
- PLATO-2: Towards Building an Open-Domain Chatbot via Curriculum Learning
- Deep Posterior Distribution-based Embedding for Hyperspectral Image Super-resolution
- SCANet: Scene Complexity Aware Network for Weakly-Supervised Video Moment Retrieval
- Interpreting Deep Learning Models in Natural Language Processing: A Review
- Developmentally motivated emergence of compositional communication via template transfer
- Instance-aware, Context-focused, and Memory-efficient Weakly Supervised Object Detection
- Hierarchical Quantized Autoencoders
- Advances in Variational Inference
- Explanation as a Defense of Recommendation
- Enhancing Neural Architecture Search with Multiple Hardware Constraints for Deep Learning Model Deployment on Tiny IoT Devices
- AdaBERT: Task-Adaptive BERT Compression with Differentiable Neural Architecture Search
- Data Imputation with Iterative Graph Reconstruction
- Unsupervised Cipher Cracking Using Discrete GANs
- Well-calibrated Model Uncertainty with Temperature Scaling for Dropout Variational Inference
- ME-D2N: Multi-Expert Domain Decompositional Network for Cross-Domain Few-Shot Learning
- Multi-Features Guidance Network for partial-to-partial point cloud registration
- Deep Transformers with Latent Depth
- Gradient Estimation with Stochastic Softmax Tricks
- Point Cloud Registration using Representative Overlapping Points
- Learning to Select Knowledge for Response Generation in Dialog Systems
- Improving Sequence-to-Sequence Learning via Optimal Transport
- Parameter-Efficient Transfer Learning with Diff Pruning
- Learning Graph Structures with Transformer for Multivariate Time Series Anomaly Detection in IoT
- Unsupervised Discrete Sentence Representation Learning for Interpretable Neural Dialog Generation
- Path Planning using Neural A* Search
- Efficient Learning of the Parameters of Non-Linear Models using Differentiable Resampling in Particle Filters
- Poincaré Wasserstein Autoencoder
- Model-Based Planning with Discrete and Continuous Actions
- Rethinking Action Spaces for Reinforcement Learning in End-to-end Dialog Agents with Latent Variable Models
- Learning Sparse Interaction Graphs of Partially Detected Pedestrians for Trajectory Prediction
- Intensity-Free Learning of Temporal Point Processes
- Scalable Rule-Based Representation Learning for Interpretable Classification
- Learning Interpretable Rules for Scalable Data Representation and Classification
- Rules of the Road: Predicting Driving Behavior with a Convolutional Model of Semantic Interactions
- not-MIWAE: Deep Generative Modelling with Missing not at Random Data
- Discrete Graph Structure Learning for Forecasting Multiple Time Series
- QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object Detection
- Structure-Aware DropEdge Towards Deep Graph Convolutional Networks
- Unsupervised Speech Recognition
- AdaFuse: Adaptive Temporal Fusion Network for Efficient Action Recognition
- Revisiting Reweighted Wake-Sleep for Models with Stochastic Control Flow
- Constrained Learning with Non-Convex Losses
- Sample-Efficient Imitation Learning via Generative Adversarial Nets
- Probabilistic Dual Network Architecture Search on Graphs
- Multi-Agent Game Abstraction via Graph Attention Neural Network
- State Space Advanced Fuzzy Cognitive Map approach for automatic and non Invasive diagnosis of Coronary Artery Disease
- Variational Memory Addressing in Generative Models
- Weakly-Supervised 3D Medical Image Segmentation using Geometric Prior and Contrastive Similarity
- TG-GAN: Continuous-time Temporal Graph Generation with Deep Generative Models
- Learnable Bernoulli Dropout for Bayesian Deep Learning
- Multi-agent Deep Reinforcement Learning with Extremely Noisy Observations
- Action Sequence Augmentation for Early Graph-based Anomaly Detection
- Differentiable Model Compression via Pseudo Quantization Noise
- MALA: Cross-Domain Dialogue Generation with Action Learning
- Model Inversion Networks for Model-Based Optimization
- ToM2C: Target-oriented Multi-agent Communication and Cooperation with Theory of Mind
- Learning compositionally through attentive guidance
- Plug & Play Directed Evolution of Proteins with Gradient-based Discrete MCMC
- Supervised and Self-supervised Pretraining Based COVID-19 Detection Using Acoustic Breathing/Cough/Speech Signals
- Memory-efficient Embedding for Recommendations
- Adversarial Contrastive Estimation
- Stochastic Beams and Where to Find Them: The Gumbel-Top-k Trick for Sampling Sequences Without Replacement
- Weight-dependent Gates for Network Pruning
- SegBlocks: Block-Based Dynamic Resolution Networks for Real-Time Segmentation
- Learning to Collocate Neural Modules for Image Captioning
- Really Useful Synthetic Data -- A Framework to Evaluate the Quality of Differentially Private Synthetic Data
- -ARM: Network Sparsification via Stochastic Binary Optimization
- HardCoRe-NAS: Hard Constrained diffeRentiable Neural Architecture Search
- Delay-Aware Multi-Agent Reinforcement Learning for Cooperative and Competitive Environments
- Robust Lottery Tickets for Pre-trained Language Models
- A Statistical Framework of Watermarks for Large Language Models: Pivot, Detection Efficiency and Optimal Rules
- Bayesian Graph Neural Networks with Adaptive Connection Sampling
- Gumbel-Attention for Multi-modal Machine Translation
- Auto-Encoding Knockoff Generator for FDR Controlled Variable Selection
- SummAE: Zero-Shot Abstractive Text Summarization using Length-Agnostic Auto-Encoders
- Estimate and Replace: A Novel Approach to Integrating Deep Neural Networks with Existing Applications
- Efficient Transformers with Dynamic Token Pooling
- Differentiable Causal Discovery from Interventional Data
- vGraph: A Generative Model for Joint Community Detection and Node Representation Learning
- Relaxed Quantization for Discretized Neural Networks
- Disentangled Recurrent Wasserstein Autoencoder
- A Review of Learning with Deep Generative Models from Perspective of Graphical Modeling
- Generating Diverse and Meaningful Captions
- Pushing Paraphrase Away from Original Sentence: A Multi-Round Paraphrase Generation Approach
- Training Binary Neural Networks using the Bayesian Learning Rule
- Estimating Gradients for Discrete Random Variables by Sampling without Replacement
- Bayesian Semisupervised Learning with Deep Generative Models
- DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task Learning
- DNA: Differentiable Network-Accelerator Co-Search
- Can Subnetwork Structure be the Key to Out-of-Distribution Generalization?
- Differentiable Top-k Operator with Optimal Transport
- An Explicit Local and Global Representation Disentanglement Framework with Applications in Deep Clustering and Unsupervised Object Detection
- Motion Planner Augmented Reinforcement Learning for Robot Manipulation in Obstructed Environments
- Learning Neural-Symbolic Descriptive Planning Models via Cube-Space Priors: The Voyage Home (to STRIPS)
- Texar: A Modularized, Versatile, and Extensible Toolkit for Text Generation
- SALSA-TEXT : self attentive latent space based adversarial text generation
- GraphGLOW: Universal and Generalizable Structure Learning for Graph Neural Networks
- Exploring Sparsity in Image Super-Resolution for Efficient Inference
- Efficient Neural Network Training via Forward and Backward Propagation Sparsification
- PRSeg: A Lightweight Patch Rotate MLP Decoder for Semantic Segmentation
- Learning to Drop: Robust Graph Neural Network via Topological Denoising
- Object Files and Schemata: Factorizing Declarative and Procedural Knowledge in Dynamical Systems
- Dirichlet Variational Autoencoder for Text Modeling
- Promoting Coordination through Policy Regularization in Multi-Agent Deep Reinforcement Learning
- PixelVAE++: Improved PixelVAE with Discrete Prior
- Elastic-InfoGAN: Unsupervised Disentangled Representation Learning in Class-Imbalanced Data
- Quantum Generative Models for Small Molecule Drug Discovery
- Synthetic Observational Health Data with GANs: from slow adoption to a boom in medical research and ultimately digital twins?
- Latent Tree Learning with Differentiable Parsers: Shift-Reduce Parsing and Chart Parsing
- Benchmarking Multi-Agent Deep Reinforcement Learning Algorithms in Cooperative Tasks
- Efficient Differentiable Neural Architecture Search with Meta Kernels
- Differentiable Unsupervised Feature Selection based on a Gated Laplacian
- Structured Reordering for Modeling Latent Alignments in Sequence Transduction
- Differentiable Predictions for Large Scale Structure with SHAMNet
- ASFD: Automatic and Scalable Face Detector
- Explaining a black-box using Deep Variational Information Bottleneck Approach
- Efficient Bitwidth Search for Practical Mixed Precision Neural Network
- Class-Weighted Classification: Trade-offs and Robust Approaches
- A Novel Attribute Reconstruction Attack in Federated Learning
- TED: A Pretrained Unsupervised Summarization Model with Theme Modeling and Denoising
- TraDE: Transformers for Density Estimation
- Densely Connected Search Space for More Flexible Neural Architecture Search
- Self-training and Pre-training are Complementary for Speech Recognition
- Dual-interest Factorization-heads Attention for Sequential Recommendation
- Theoretical Insights Into Multiclass Classification: A High-dimensional Asymptotic View
- Learning Navigation Subroutines from Egocentric Videos
- Generating Natural Language Adversarial Examples on a Large Scale with Generative Models
- Graphical Normalizing Flows
- CMDNet: Learning a Probabilistic Relaxation of Discrete Variables for Soft Detection with Low Complexity
- LAP-Net: Adaptive Features Sampling via Learning Action Progression for Online Action Detection
- Fast, Diverse and Accurate Image Captioning Guided By Part-of-Speech
- Dual Attention Networks for Visual Reference Resolution in Visual Dialog
- Adversarial Example Games
- Guided Dialog Policy Learning without Adversarial Learning in the Loop
- Information Leakage in Embedding Models
- Deep Generation of Heterogeneous Networks
- End-to-End Entity Linking and Disambiguation leveraging Word and Knowledge Graph Embeddings
- FBNetV5: Neural Architecture Search for Multiple Tasks in One Run
- Learning to simulate and design for structural engineering
- Learning to Hash with Graph Neural Networks for Recommender Systems
- A Simple Probabilistic Method for Deep Classification under Input-Dependent Label Noise
- A Comprehensive Survey of Deep Learning for Image Captioning
- Improving the quality of generative models through Smirnov transformation
- Layer Adaptive Node Selection in Bayesian Neural Networks: Statistical Guarantees and Implementation Details
- Controlling Computation versus Quality for Neural Sequence Models
- Unsupervised Dialog Structure Learning
- The continuous categorical: a novel simplex-valued exponential family
- Creative GANs for generating poems, lyrics, and metaphors
- Categorical EHR Imputation with Generative Adversarial Nets
- Bayesian Learning of Neural Network Architectures
- Emergent Communication Pretraining for Few-Shot Machine Translation
- R2D2: Recursive Transformer based on Differentiable Tree for Interpretable Hierarchical Language Modeling
- Learning K-way D-dimensional Discrete Code For Compact Embedding Representations
- Pairwise Supervised Hashing with Bernoulli Variational Auto-Encoder and Self-Control Gradient Estimator
- Vector representations of text data in deep learning
- Private Post-GAN Boosting
- Semi-Supervised Generation with Cluster-aware Generative Models
- Structural Inference of Networked Dynamical Systems with Universal Differential Equations
- Differentiable Architecture Search with Ensemble Gumbel-Softmax
- Visual-Semantic Transformer for Scene Text Recognition
- CompILE: Compositional Imitation Learning and Execution
- Universal Adversarial Attacks with Natural Triggers for Text Classification
- Image-Question-Answer Synergistic Network for Visual Dialog
- Calibration of Model Uncertainty for Dropout Variational Inference
- Topic-Driven and Knowledge-Aware Transformer for Dialogue Emotion Detection
- Improving GAN Training with Probability Ratio Clipping and Sample Reweighting
- Adversarial Learned Molecular Graph Inference and Generation
- Efficient Marginalization of Discrete and Structured Latent Variables via Sparsity
- Simple and Effective Zero-shot Cross-lingual Phoneme Recognition
- Learning from Lexical Perturbations for Consistent Visual Question Answering
- Spatially Adaptive Inference with Stochastic Feature Sampling and Interpolation
- Real-time Denoising and Dereverberation with Tiny Recurrent U-Net
- Modelling Latent Translations for Cross-Lingual Transfer
- Learn to Explain Efficiently via Neural Logic Inductive Learning
- Counterfactual Explanations for Arbitrary Regression Models
- Elephant in the Room: An Evaluation Framework for Assessing Adversarial Examples in NLP
- Variable-rate discrete representation learning
- Interpretable and Pedagogical Examples
- Cooperative Multi-Agent Transfer Learning with Level-Adaptive Credit Assignment
- Composing Task-Agnostic Policies with Deep Reinforcement Learning
- Efficient Neural Architecture Search via Proximal Iterations
- Learning Compressed Sentence Representations for On-Device Text Processing
- Networked Multi-Agent Reinforcement Learning with Emergent Communication
- Efficient Deep Neural Networks
- AR-Net: Adaptive Frame Resolution for Efficient Action Recognition
- Residual Correction in Real-Time Traffic Forecasting
- vONTSS: vMF based semi-supervised neural topic modeling with optimal transport
- Exploring TTS without T Using Biologically/Psychologically Motivated Neural Network Modules (ZeroSpeech 2020)
- Neural Sparse Representation for Image Restoration
- Learning Visual Question Answering by Bootstrapping Hard Attention
- The Emergence of Wireless MAC Protocols with Multi-Agent Reinforcement Learning
- Interpretable agent communication from scratch (with a generic visual processor emerging on the side)
- Turn the Combination Lock: Learnable Textual Backdoor Attacks via Word Substitution
- Contrastive Learning with Adversarial Perturbations for Conditional Text Generation
- Multi-space Variational Encoder-Decoders for Semi-supervised Labeled Sequence Transduction
- Sensor-Augmented Egocentric-Video Captioning with Dynamic Modal Attention
- Dynamic Frame Interpolation in Wavelet Domain
- Rationalizing Predictions by Adversarial Information Calibration
- Differentiable Logic Machines
- Inferring Multidimensional Rates of Aging from Cross-Sectional Data
- SUMBT+LaRL: Effective Multi-domain End-to-end Neural Task-oriented Dialog System
- DisARM: An Antithetic Gradient Estimator for Binary Latent Variables
- Improving Generative Imagination in Object-Centric World Models
- Towards Diverse Paraphrase Generation Using Multi-Class Wasserstein GAN
- Reparameterization Gradient for Non-differentiable Models
- BERE: An accurate distantly supervised biomedical entity relation extraction network
- ChartPointFlow for Topology-Aware 3D Point Cloud Generation
- InfoCatVAE: Representation Learning with Categorical Variational Autoencoders
- Learning Graph-Level Representations with Recurrent Neural Networks
- Riemannian Normalizing Flow on Variational Wasserstein Autoencoder for Text Modeling
- Information Maximizing Visual Question Generation
- A Latent Morphology Model for Open-Vocabulary Neural Machine Translation
- Emergent Communication of Generalizations
- Using Large Ensembles of Control Variates for Variational Inference
- VX2TEXT: End-to-End Learning of Video-Based Text Generation From Multimodal Inputs
- Drop-Bottleneck: Learning Discrete Compressed Representation for Noise-Robust Exploration
- Towards Improving the Consistency, Efficiency, and Flexibility of Differentiable Neural Architecture Search
- ACtuAL: Actor-Critic Under Adversarial Learning
- NASI: Label- and Data-agnostic Neural Architecture Search at Initialization
- Towards Federated Bayesian Network Structure Learning with Continuous Optimization
- Text Summarization with Latent Queries
- Deriving Machine Attention from Human Rationales
- Learning Hierarchical Teaching Policies for Cooperative Agents
- Semantic Bottleneck Scene Generation
- Topic-Preserving Synthetic News Generation: An Adversarial Deep Reinforcement Learning Approach
- Unsupervised Domain Adaptation for Semantic Segmentation via Low-level Edge Information Transfer
- In the Eye of the Beholder: Gaze and Actions in First Person Video
- Modeling Latent Sentence Structure in Neural Machine Translation
- Learning Discrete Structured Representations by Adversarially Maximizing Mutual Information
- Aggregate or Not? Exploring Where to Privatize in DNN Based Federated Learning Under Different Non-IID Scenes
- ENGINE: Energy-Based Inference Networks for Non-Autoregressive Machine Translation
- Automated Model Design and Benchmarking of 3D Deep Learning Models for COVID-19 Detection with Chest CT Scans
- SSN: Learning Sparse Switchable Normalization via SparsestMax
- Radial and Directional Posteriors for Bayesian Neural Networks
- Dynamic Compositional Graph Convolutional Network for Efficient Composite Human Motion Prediction
- GAEA: Graph Augmentation for Equitable Access via Reinforcement Learning
- Storchastic: A Framework for General Stochastic Automatic Differentiation
- Disentanglement via Mechanism Sparsity Regularization: A New Principle for Nonlinear ICA
- Cascaded channel pruning using hierarchical self-distillation
- On Sparsifying Encoder Outputs in Sequence-to-Sequence Models
- Soft then Hard: Rethinking the Quantization in Neural Image Compression
- Recognizing Predictive Substructures with Subgraph Information Bottleneck
- Point-less: More Abstractive Summarization with Pointer-Generator Networks
- RelEx: A Model-Agnostic Relational Model Explainer
- Mask-GVAE: Blind Denoising Graphs via Partition
- Faster Meta Update Strategy for Noise-Robust Deep Learning
- Learning a Multi-Modal Policy via Imitating Demonstrations with Mixed Behaviors
- CAT-Gen: Improving Robustness in NLP Models via Controlled Adversarial Text Generation
- Gaussian mixture models with Wasserstein distance
- Learning to Ask Questions in Open-domain Conversational Systems with Typed Decoders
- Forecasting Human-Object Interaction: Joint Prediction of Motor Attention and Actions in First Person Video
- PETGEN: Personalized Text Generation Attack on Deep Sequence Embedding-based Classification Models
- Maximum Mean Discrepancy for Generalization in the Presence of Distribution and Missingness Shift
- Conditional Generative Models for Counterfactual Explanations
- The Bayesian Learning Rule
- Hierarchical VampPrior Variational Fair Auto-Encoder
- FairIF: Boosting Fairness in Deep Learning via Influence Functions with Validation Set Sensitive Attributes
- Image Captioning Based on a Hierarchical Attention Mechanism and Policy Gradient Optimization
- Deep Learning for Click-Through Rate Estimation
- Have We Learned to Explain?: How Interpretability Methods Can Learn to Encode Predictions in their Interpretations
- Effective and Fast: A Novel Sequential Single Path Search for Mixed-Precision Quantization
- Entity-Consistent End-to-end Task-Oriented Dialogue System with KB Retriever
- Predictive Sampling with Forecasting Autoregressive Models
- Bayesian Inference on Binary Spiking Networks Leveraging Nanoscale Device Stochasticity
- Emergent Discrete Communication in Semantic Spaces
- Object-Centric Image Generation with Factored Depths, Locations, and Appearances
- Towards Interpretable and Reliable Reading Comprehension: A Pipeline Model with Unanswerability Prediction
- Lifelong Bayesian Optimization
- Non-Differentiable Supervised Learning with Evolution Strategies and Hybrid Methods
- Generating Multiple Diverse Responses with Multi-Mapping and Posterior Mapping Selection
- Learning to Discretely Compose Reasoning Module Networks for Video Captioning
- SoftSort: A Continuous Relaxation for the argsort Operator
- Phase-aware Single-stage Speech Denoising and Dereverberation with U-Net
- Discovering Dialog Structure Graph for Open-Domain Dialog Generation
- Implicit MLE: Backpropagating Through Discrete Exponential Family Distributions
- SpVOS: Efficient Video Object Segmentation with Triple Sparse Convolution
- Differentiable Retrieval Augmentation via Generative Language Modeling for E-commerce Query Intent Classification
- Cross-Modal Conceptualization in Bottleneck Models
- Accelerating SNN Training with Stochastic Parallelizable Spiking Neurons
- Unleashing Transformers: Parallel Token Prediction with Discrete Absorbing Diffusion for Fast High-Resolution Image Generation from Vector-Quantized Codes
- SQUID: Deep Feature In-Painting for Unsupervised Anomaly Detection
- Towards Empirical Sandwich Bounds on the Rate-Distortion Function
- Learning Frequency-aware Dynamic Network for Efficient Super-Resolution
- AMEIR: Automatic Behavior Modeling, Interaction Exploration and MLP Investigation in the Recommender System
- Fast and Flexible Temporal Point Processes with Triangular Maps
- Recommender Systems Based on Generative Adversarial Networks: A Problem-Driven Perspective
- Sparse Graph Attention Networks
- Particle Smoothing Variational Objectives
- Mixture Content Selection for Diverse Sequence Generation
- Amortized Bethe Free Energy Minimization for Learning MRFs
- Neural Clustering Processes
- Modeling Attention Flow on Graphs
- Hierarchical Approaches for Reinforcement Learning in Parameterized Action Space
- LeMoNADe: Learned Motif and Neuronal Assembly Detection in calcium imaging videos
- Do latent tree learning models identify meaningful structure in sentences?
- Scalable Deep Unsupervised Clustering with Concrete GMVAEs
- Tracking by Animation: Unsupervised Learning of Multi-Object Attentive Trackers
- Direct Optimization through for Discrete Variational Auto-Encoder
- CCGG: A Deep Autoregressive Model for Class-Conditional Graph Generation
- Know What You Don't Need: Single-Shot Meta-Pruning for Attention Heads
- CatGAN: Category-aware Generative Adversarial Networks with Hierarchical Evolutionary Learning for Category Text Generation
- Deep clustering with concrete k-means
- Weakly Supervised Explainable Phrasal Reasoning with Neural Fuzzy Logic
- UFO-BLO: Unbiased First-Order Bilevel Optimization
- Mastering emergent language: learning to guide in simulated navigation
- Pay Attention when Required
- Computation on Sparse Neural Networks: an Inspiration for Future Hardware
- Multi-Facet Clustering Variational Autoencoders
- The Out-of-Distribution Problem in Explainability and Search Methods for Feature Importance Explanations
- BARS: Joint Search of Cell Topology and Layout for Accurate and Efficient Binary ARchitectures
- Generative Hybrid Representations for Activity Forecasting with No-Regret Learning
- Deep Residual Mixture Models
- In-Distribution Interpretability for Challenging Modalities
- Learning with Algorithmic Supervision via Continuous Relaxations
- Self-Supervised Bernoulli Autoencoders for Semi-Supervised Hashing
- Pay Better Attention to Attention: Head Selection in Multilingual and Multi-Domain Sequence Modeling
- Failure Modes of Variational Autoencoders and Their Effects on Downstream Tasks
- Towards Stable Symbol Grounding with Zero-Suppressed State AutoEncoder
- AutoLoss: Automated Loss Function Search in Recommendations
- Hierarchical Multi-scale Attention Networks for Action Recognition
- Hybrid Generative-Contrastive Representation Learning
- SHOT-VAE: Semi-supervised Deep Generative Models With Label-aware ELBO Approximations
- Judge the Judges: A Large-Scale Evaluation Study of Neural Language Models for Online Review Generation
- Are Training Resources Insufficient? Predict First Then Explain!
- Product Kanerva Machines: Factorized Bayesian Memory
- Learning Architectures for Binary Networks
- Leap-LSTM: Enhancing Long Short-Term Memory for Text Categorization
- Latent Programmer: Discrete Latent Codes for Program Synthesis
- Recursive Visual Attention in Visual Dialog
- Discovering Predictive Relational Object Symbols with Symbolic Attentive Layers
- End-to-End Feedback Loss in Speech Chain Framework via Straight-Through Estimator
- Learning Variational Word Masks to Improve the Interpretability of Neural Text Classifiers
- A Unified Pre-training Framework for Conversational AI
- Environment Invariant Linear Least Squares
- Latent Iterative Refinement for Modular Source Separation
- KinePose: A temporally optimized inverse kinematics technique for 6DOF human pose estimation with biomechanical constraints
- Improved Gradient-Based Optimization Over Discrete Distributions
- Joint Visual Semantic Reasoning: Multi-Stage Decoder for Text Recognition
- Perceive, Attend, and Drive: Learning Spatial Attention for Safe Self-Driving
- Mitigating Gender Bias for Neural Dialogue Generation with Adversarial Learning
- Benchmarking Unsupervised Object Representations for Video Sequences
- A Biologically Inspired Visual Working Memory for Deep Networks
- Cross-Attention Watermarking of Large Language Models
- Searching for Stage-wise Neural Graphs In the Limit
- Bayesian Structure Adaptation for Continual Learning
- An Empirical Study: Extensive Deep Temporal Point Process
- Sparse Stochastic Zeroth-Order Optimization with an Application to Bandit Structured Prediction
- RepNAS: Searching for Efficient Re-parameterizing Blocks
- Index Tracking with Cardinality Constraints: A Stochastic Neural Networks Approach
- On Machine Learning and Structure for Mobile Robots
- Triggering Dark Showers with Conditional Dual Auto-Encoders
- Avoiding hashing and encouraging visual semantics in referential emergent language games
- Collapsed Amortized Variational Inference for Switching Nonlinear Dynamical Systems
- Learning Multi-Task Transferable Rewards via Variational Inverse Reinforcement Learning
- MALCOM: Generating Malicious Comments to Attack Neural Fake News Detection Models
- Unsupervised Action Segmentation for Instructional Videos
- Deep Learning from Noisy Image Labels with Quality Embedding
- Caching Transient Content for IoT Sensing: Multi-Agent Soft Actor-Critic
- MetaMix: Meta-state Precision Searcher for Mixed-precision Activation Quantization
- When in Doubt: Neural Non-Parametric Uncertainty Quantification for Epidemic Forecasting
- Towards Emergent Language Symbolic Semantic Segmentation and Model Interpretability
- EASE: Extractive-Abstractive Summarization with Explanations
- AutoFT: Automatic Fine-Tune for Parameters Transfer Learning in Click-Through Rate Prediction
- : Author Attribute Anonymity by Adversarial Training of Neural Machine Translation
- Deep Variational Transfer: Transfer Learning through Semi-supervised Deep Generative Models
- Effective Sparsification of Neural Networks with Global Sparsity Constraint
- InfoFair: Information-Theoretic Intersectional Fairness
- To Relieve Your Headache of Training an MRF, Take AdVIL
- Learning to learn generative programs with Memoised Wake-Sleep
- Unsupervised and interpretable scene discovery with Discrete-Attend-Infer-Repeat
- Learning to Compose Hypercolumns for Visual Correspondence
- User-specific Adaptive Fine-tuning for Cross-domain Recommendations
- RetrieveGAN: Image Synthesis via Differentiable Patch Retrieval
- The Referential Reader: A Recurrent Entity Network for Anaphora Resolution
- Fine-Grained Stochastic Architecture Search
- Latent Domain Learning with Dynamic Residual Adapters
- Entropy Minimization In Emergent Languages
- OOGAN: Disentangling GAN with One-Hot Sampling and Orthogonal Regularization
- The Thermodynamic Variational Objective
- HorNet: A Hierarchical Offshoot Recurrent Network for Improving Person Re-ID via Image Captioning
- PONAS: Progressive One-shot Neural Architecture Search for Very Efficient Deployment
- Cooperative Learning of Disjoint Syntax and Semantics
- Disentangling to Cluster: Gaussian Mixture Variational Ladder Autoencoders
- Towards Non-saturating Recurrent Units for Modelling Long-term Dependencies
- Improving Disentangled Representation Learning with the Beta Bernoulli Process
- What do Deep Networks Like to Read?
- GO Gradient for Expectation-Based Objectives
- Adversarial Training for Community Question Answer Selection Based on Multi-scale Matching
- Grey-box Adversarial Attack And Defence For Sentiment Classification
- Countering Language Drift with Seeded Iterated Learning
- Neural Generators of Sparse Local Linear Models for Achieving both Accuracy and Interpretability
- Contextual Temperature for Language Modeling
- Learning from the Best: Rationalizing Prediction by Adversarial Information Calibration
- A Study of Joint Graph Inference and Forecasting
- Weakly Supervised Reasoning by Neuro-Symbolic Approaches
- Inductive Bias and Language Expressivity in Emergent Communication
- A Probabilistic End-To-End Task-Oriented Dialog Model with Latent Belief States towards Semi-Supervised Learning
- Unsupervised Transfer of Semantic Role Models from Verbal to Nominal Domain
- Improved Variational Neural Machine Translation by Promoting Mutual Information
- Semi-Supervised Variational Autoencoder for Survival Prediction
- Decomposable Neural Paraphrase Generation
- StampNet: unsupervised multi-class object discovery
- Bi-level Score Matching for Learning Energy-based Latent Variable Models
- Variational Inference with Holder Bounds
- Safeguarded Dynamic Label Regression for Generalized Noisy Supervision
- DS-Net++: Dynamic Weight Slicing for Efficient Inference in CNNs and Transformers
- IOT: Instance-wise Layer Reordering for Transformer Structures
- MoBoAligner: a Neural Alignment Model for Non-autoregressive TTS with Monotonic Boundary Search
- Learning Synthetic Environments for Reinforcement Learning with Evolution Strategies
- Correcting Experience Replay for Multi-Agent Communication
- CoDiNet: Path Distribution Modeling with Consistency and Diversity for Dynamic Routing
- Boosting Naturalness of Language in Task-oriented Dialogues via Adversarial Training
- Variance Reduction for Evolution Strategies via Structured Control Variates
- Discond-VAE: Disentangling Continuous Factors from the Discrete
- Drug-Drug Interaction Prediction with Wasserstein Adversarial Autoencoder-based Knowledge Graph Embeddings
- ASL Recognition with Metric-Learning based Lightweight Network
- Dissimilarity Mixture Autoencoder for Deep Clustering
- Backprop-Q: Generalized Backpropagation for Stochastic Computation Graphs
- A perspective on multi-agent communication for information fusion
- Emergence of Numeric Concepts in Multi-Agent Autonomous Communication
- Generative Hierarchical Models for Parts, Objects, and Scenes
- Effect of latent space distribution on the segmentation of images with multiple annotations
- Weakly Supervised Concept Map Generation through Task-Guided Graph Translation
- Semi-supervised Disentanglement with Independent Vector Variational Autoencoders
- Neural Execution of Graph Algorithms
- Unsupervised Text Embedding Space Generation Using Generative Adversarial Networks for Text Synthesis
- Exploiting Operation Importance for Differentiable Neural Architecture Search
- LAVAE: Disentangling Location and Appearance
- Examining the Ordering of Rhetorical Strategies in Persuasive Requests
- Rephrasing visual questions by specifying the entropy of the answer distribution
- Rethinking Supervised Learning and Reinforcement Learning in Task-Oriented Dialogue Systems
- Coupled Gradient Estimators for Discrete Latent Variables
- Is Bang-Bang Control All You Need? Solving Continuous Control with Bernoulli Policies
- Conditional Hybrid GAN for Sequence Generation
- Inducing Grammars with and for Neural Machine Translation
- EMAP: Explanation by Minimal Adversarial Perturbation
- Adding A Filter Based on The Discriminator to Improve Unconditional Text Generation
- Inducing and Embedding Senses with Scaled Gumbel Softmax
- FEAR: A Simple Lightweight Method to Rank Architectures
- Transformer-Based Local Feature Matching for Multimodal Image Registration
- Sampling-Free Learning of Bayesian Quantized Neural Networks
- Sentence Encoding with Tree-constrained Relation Networks
- Data Manipulation: Towards Effective Instance Learning for Neural Dialogue Generation via Learning to Augment and Reweight
- Neural-Symbolic Descriptive Action Model from Images: The Search for STRIPS
- Stochastic Sequential Neural Networks with Structured Inference
- Path Sample-Analytic Gradient Estimators for Stochastic Binary Networks
- Auto-Agent-Distiller: Towards Efficient Deep Reinforcement Learning Agents via Neural Architecture Search
- Transferable Time-Series Forecasting under Causal Conditional Shift
- High-Resolution Complex Scene Synthesis with Transformers
- High-Capacity Expert Binary Networks
- Discrete-continuous Action Space Policy Gradient-based Attention for Image-Text Matching
- Variational Rejection Sampling
- An Empirical Investigation of Global and Local Normalization for Recurrent Neural Sequence Models Using a Continuous Relaxation to Beam Search
- Evidential Sparsification of Multimodal Latent Spaces in Conditional Variational Autoencoders
- VQ-DRAW: A Sequential Discrete VAE
- Efficient Neural Architecture Search for End-to-end Speech Recognition via Straight-Through Gradients
- Continual Learning: Tackling Catastrophic Forgetting in Deep Neural Networks with Replay Processes
- GRADE: Graph Dynamic Embedding
- Probabilistic Mixture-of-Experts for Efficient Deep Reinforcement Learning
- Differentiable Equilibrium Computation with Decision Diagrams for Stackelberg Models of Combinatorial Congestion Games
- Learning Randomly Perturbed Structured Predictors for Direct Loss Minimization
- Dynamic Network Quantization for Efficient Video Inference
- Multi-sense Definition Modeling using Word Sense Decompositions
- Population-based Gradient Descent Weight Learning for Graph Coloring Problems
- Combining Two Adversarial Attacks Against Person Re-Identification Systems
- Adversarial Sub-sequence for Text Generation
- Controlling Text Edition by Changing Answers of Specific Questions
- An Imitation Learning Approach to Unsupervised Parsing
- NomMer: Nominate Synergistic Context in Vision Transformer for Visual Recognition
- Variational Latent-State GPT for Semi-Supervised Task-Oriented Dialog Systems
- Variational Marginal Particle Filters
- Learning Slice-Aware Representations with Mixture of Attentions
- Auto-NBA: Efficient and Effective Search Over the Joint Space of Networks, Bitwidths, and Accelerators
- Expressivity of Emergent Language is a Trade-off between Contextual Complexity and Unpredictability
- NWT: Towards natural audio-to-video generation with representation learning
- Discriminative Triad Matching and Reconstruction for Weakly Referring Expression Grounding
- BWCP: Probabilistic Learning-to-Prune Channels for ConvNets via Batch Whitening
- Connecting What to Say With Where to Look by Modeling Human Attention Traces
- Iterated learning for emergent systematicity in VQA
- Text Generation with Deep Variational GAN
- Gradient-based Adversarial Attacks against Text Transformers
- Direct Differentiable Augmentation Search
- NewsBERT: Distilling Pre-trained Language Model for Intelligent News Application
- Enhancing Content Preservation in Text Style Transfer Using Reverse Attention and Conditional Layer Normalization
- A Framework for Joint Unsupervised Learning of Cluster-Aware Embedding for Heterogeneous Networks
- Causal Attention for Unbiased Visual Recognition
- Investigating and Simplifying Masking-based Saliency Methods for Model Interpretability
- Dual Reconstruction: a Unifying Objective for Semi-Supervised Neural Machine Translation
- On-device neural speech synthesis
- Optimal Variance Control of the Score Function Gradient Estimator for Importance Weighted Bounds
- Integrated Training for Sequence-to-Sequence Models Using Non-Autoregressive Transformer
- Graph-Based Continual Learning
- Concept-Aware Denoising Graph Neural Network for Micro-Video Recommendation
- BNAS v2: Learning Architectures for Binary Networks with Empirical Improvements
- FP-NAS: Fast Probabilistic Neural Architecture Search
- Exploring Contextual Word-level Style Relevance for Unsupervised Style Transfer
- Self-supervised Segmentation via Background Inpainting
- Generating Semantically Valid Adversarial Questions for TableQA
- Uncertainty Estimation and Calibration with Finite-State Probabilistic RNNs
- Writing Polishment with Simile: Task, Dataset and A Neural Approach
- On (Emergent) Systematic Generalisation and Compositionality in Visual Referential Games with Straight-Through Gumbel-Softmax Estimator
- Meta-CoTGAN: A Meta Cooperative Training Paradigm for Improving Adversarial Text Generation
- GANs with Conditional Independence Graphs: On Subadditivity of Probability Divergences
- Weakly-Supervised Hierarchical Models for Predicting Persuasive Strategies in Good-faith Textual Requests
- Structural Inductive Biases in Emergent Communication
- Neural Data-to-Text Generation with LM-based Text Augmentation
- Do as I mean, not as I say: Sequence Loss Training for Spoken Language Understanding
- GLAM: Graph Learning by Modeling Affinity to Labeled Nodes for Graph Neural Networks
- Invertible Gaussian Reparameterization: Revisiting the Gumbel-Softmax
- Alleviate Exposure Bias in Sequence Prediction \\ with Recurrent Neural Networks
- Adaptive Correlated Monte Carlo for Contextual Categorical Sequence Generation
- Prototype-based Personalized Pruning
- Imperfect also Deserves Reward: Multi-Level and Sequential Reward Modeling for Better Dialog Management
- Explaining Neural Network Predictions on Sentence Pairs via Learning Word-Group Masks
- An Auxiliary Classifier Generative Adversarial Framework for Relation Extraction
- Bi-level Actor-Critic for Multi-agent Coordination
- Feature Gradients: Scalable Feature Selection via Discrete Relaxation
- TimeGate: Conditional Gating of Segments in Long-range Activities
- Neural Image Compression and Explanation
- An Unsupervised Bayesian Neural Network for Truth Discovery in Social Networks
- Dispersed Exponential Family Mixture VAEs for Interpretable Text Generation
- Neural Stored-program Memory
- Community Detection Clustering via Gumbel Softmax
- Learning Robust Feature Representations for Scene Text Detection
- Learning to refer informatively by amortizing pragmatic reasoning
- A Meta-Bayesian Model of Intentional Visual Search
- Semi-supervised Sequential Generative Models
- Question Guided Modular Routing Networks for Visual Question Answering
- Plug-in, Trainable Gate for Streamlining Arbitrary Neural Networks
- Winning an Election: On Emergent Strategic Communication in Multi-Agent Networks
- Foveation for Segmentation of Ultra-High Resolution Images
- OCEAN: Online Task Inference for Compositional Tasks with Context Adaptation
- Semi-Supervised Confidence Network aided Gated Attention based Recurrent Neural Network for Clickbait Detection
- Differentiable Greedy Networks
- Learning to Forecast Videos of Human Activity with Multi-granularity Models and Adaptive Rendering
- Energy-Inspired Models: Learning with Sampler-Induced Distributions
- A unified view of likelihood ratio and reparameterization gradients and an optimal importance sampling scheme
- Discrete Structural Planning for Neural Machine Translation
- Learning Product Codebooks using Vector Quantized Autoencoders for Image Retrieval
- Not All Attention Is Needed: Gated Attention Network for Sequence Data
- Conformation Clustering of Long MD Protein Dynamics with an Adversarial Autoencoder
- Discrete flow posteriors for variational inference in discrete dynamical systems
- NASH: Toward End-to-End Neural Architecture for Generative Semantic Hashing
- Towards Graph Representation Learning in Emergent Communication
- Completely Unsupervised Phoneme Recognition by Adversarially Learning Mapping Relationships from Audio Embeddings
- Learning to Adaptively Scale Recurrent Neural Networks
- Actional-Structural Graph Convolutional Networks for Skeleton-based Action Recognition
- Inverting Variational Autoencoders for Improved Generative Accuracy
- Neural ODEs with stochastic vector field mixtures
- Building Minimal and Reusable Causal State Abstractions for Reinforcement Learning
- Phone and speaker spatial organization in self-supervised speech representations
- AskewSGD : An Annealed interval-constrained Optimisation method to train Quantized Neural Networks
- A Doubly Stochastic Simulator with Applications in Arrivals Modeling and Simulation
- Any-to-One Sequence-to-Sequence Voice Conversion using Self-Supervised Discrete Speech Representations
- Real-Time Semantic Segmentation via Auto Depth, Downsampling Joint Decision and Feature Aggregation
- NASA: Neural Architecture Search and Acceleration for Hardware Inspired Hybrid Networks
- S2cGAN: Semi-Supervised Training of Conditional GANs with Fewer Labels
- Context-dependent self-exciting point processes: models, methods, and risk bounds in high dimensions
- Mixture-of-Partitions: Infusing Large Biomedical Knowledge Graphs into BERT
- Multi-Aspect Temporal Network Embedding: A Mixture of Hawkes Process View
- PFPN: Continuous Control of Physically Simulated Characters using Particle Filtering Policy Network
- Math Word Problem Generation with Mathematical Consistency and Problem Context Constraints
- -Neighbor Based Curriculum Sampling for Sequence Prediction
- Joint Spatial and Layer Attention for Convolutional Networks
- Semi-Supervised Disentanglement of Class-Related and Class-Independent Factors in VAE
- Learning Generalized Gumbel-max Causal Mechanisms
- Semi-parametric Network Structure Discovery Models
- Cauchy-Schwarz Regularized Autoencoder
- Domain Decluttering: Simplifying Images to Mitigate Synthetic-Real Domain Shift and Improve Depth Estimation
- Amortised Learning by Wake-Sleep
- Deep Generative Pattern-Set Mixture Models for Nonignorable Missingness
- Pre-Learning Environment Representations for Data-Efficient Neural Instruction Following
- Focus on the Core: Efficient Attention via Pruned Token Compression for Document Classification
- Learning with Multiplicative Perturbations
- Masked Adversarial Generation for Neural Machine Translation
- Dynamic Slimmable Network
- SAG-VAE: End-to-end Joint Inference of Data Representations and Feature Relations
- Reconciling the Discrete-Continuous Divide: Towards a Mathematical Theory of Sparse Communication
- Low-Complexity Probing via Finding Subnetworks
- Adaptive Variational Bayesian Inference for Sparse Deep Neural Network
- Learning Effective and Efficient Embedding via an Adaptively-Masked Twins-based Layer
- Learning Dynamic Network Using a Reuse Gate Function in Semi-supervised Video Object Segmentation
- Neural Gaussian Copula for Variational Autoencoder
- Yet Another Representation of Binary Decision Trees: A Mathematical Demonstration
- Matching Embeddings for Domain Adaptation
- Dependent Multi-Task Learning with Causal Intervention for Image Captioning
- Neural Conditional Event Time Models
- Improving Lossless Compression Rates via Monte Carlo Bits-Back Coding
- Variational Learning for Unsupervised Knowledge Grounded Dialogs
- Off-policy Learning for Multiple Loggers
- Multilingual Zero Resource Speech Recognition Base on Self-Supervise Pre-Trained Acoustic Models
- Few-shot Learning for Unsupervised Feature Selection
- Hybrid system identification using switching density networks
- Semi-supervised Neural Chord Estimation Based on a Variational Autoencoder with Latent Chord Labels and Features
- How Does Selective Mechanism Improve Self-Attention Networks?
- BiX-NAS: Searching Efficient Bi-directional Architecture for Medical Image Segmentation
- VMI-VAE: Variational Mutual Information Maximization Framework for VAE With Discrete and Continuous Priors
- Deep Industrial Espionage
- Dynamic Routing Networks
- Learning Multiple Stock Trading Patterns with Temporal Routing Adaptor and Optimal Transport
- Implicit Integration of Superpixel Segmentation into Fully Convolutional Networks
- A Chain Graph Interpretation of Real-World Neural Networks
- Variational Mutual Information Maximization Framework for VAE Latent Codes with Continuous and Discrete Priors
- Linear, or Non-Linear, That is the Question!
- SparseGAN: Sparse Generative Adversarial Network for Text Generation
- Hierarchical Text Classification Using Contrastive Learning Informed Path Guided Hierarchy
- Wish you were here: Hindsight Goal Selection for long-horizon dexterous manipulation
- Imitation Learning of Factored Multi-agent Reactive Models
- Decentralized Multi-Agents by Imitation of a Centralized Controller
- Knowing When to Stop: Evaluation and Verification of Conformity to Output-size Specifications
- Toward Accurate and Realistic Outfits Visualization with Attention to Details
- Model-based Multi-agent Policy Optimization with Adaptive Opponent-wise Rollouts
- Learning the Solution Manifold in Optimization and Its Application in Motion Planning
- Egocentric Activity Recognition and Localization on a 3D Map
- Zero-Shot Generalization using Intrinsically Motivated Compositional Emergent Protocols
- Disentanglement of Latent Representations via Causal Interventions
- Differentiable Sampling with Flexible Reference Word Order for Neural Machine Translation
- You Never Cluster Alone
- Learning to Embed Sentences Using Attentive Recursive Trees
- Adversarial Learning of Label Dependency: A Novel Framework for Multi-class Classification
- Trainable Adaptive Window Switching for Speech Enhancement
- pix2rule: End-to-end Neuro-symbolic Rule Learning
- Iterative Refinement in the Continuous Space for Non-Autoregressive Neural Machine Translation
- Gator: Customizable Channel Pruning of Neural Networks with Gating
- Deep Switching State Space Model (DSM) for Nonlinear Time Series Forecasting with Regime Switching
- A3C-S: Automated Agent Accelerator Co-Search towards Efficient Deep Reinforcement Learning
- StyleDGPT: Stylized Response Generation with Pre-trained Language Models
- Differentiable Antithetic Sampling for Variance Reduction in Stochastic Variational Inference
- CLSRIL-23: Cross Lingual Speech Representations for Indic Languages
- Proximal Policy Optimization for Improved Convergence in IRGAN
- GMAIR: Unsupervised Object Detection Based on Spatial Attention and Gaussian Mixture
- Dynamic Resolution Network
- Fast and Efficient Scene Categorization for Autonomous Driving using VAEs
- On Tree-Based Neural Sentence Modeling
- Searching for A Robust Neural Architecture in Four GPU Hours
- Mixture factorized auto-encoder for unsupervised hierarchical deep factorization of speech signal
- Sparse Communication via Mixed Distributions
- Focus on What's Informative and Ignore What's not: Communication Strategies in a Referential Game
- Rapid Elastic Architecture Search under Specialized Classes and Resource Constraints
- LA3: Efficient Label-Aware AutoAugment
- Learning Meta Representations for Agents in Multi-Agent Reinforcement Learning
- Soft Actor-Critic With Integer Actions
- Pixel Adaptive Filtering Units
- Bounded logit attention: Learning to explain image classifiers
- Pathwise Derivatives for Multivariate Distributions
- Relaxed-Responsibility Hierarchical Discrete VAEs
- Mixture Representation Learning with Coupled Autoencoders
- Weighing Features of Lung and Heart Regions for Thoracic Disease Classification
- Protecting Anonymous Speech: A Generative Adversarial Network Methodology for Removing Stylistic Indicators in Text
- Adversarially-learned Inference via an Ensemble of Discrete Undirected Graphical Models
- Composed Fine-Tuning: Freezing Pre-Trained Denoising Autoencoders for Improved Generalization
- Optimal Operation of a Hydrogen-based Building Multi-Energy System Based on Deep Reinforcement Learning
- Discrete Word Embedding for Logical Natural Language Understanding
- CARMS: Categorical-Antithetic-REINFORCE Multi-Sample Gradient Estimator
- Greedy UnMixing for Q-Learning in Multi-Agent Reinforcement Learning
- Reliable Categorical Variational Inference with Mixture of Discrete Normalizing Flows
- Variational (Gradient) Estimate of the Score Function in Energy-based Latent Variable Models
- -based Sparse Canonical Correlation Analysis
- Integrating Human Gaze into Attention for Egocentric Activity Recognition
- Neural Latent Dependency Model for Sequence Labeling
- A New Framework for Registration of Semantic Point Clouds from Stereo and RGB-D Cameras
- Learning Discrete Energy-based Models via Auxiliary-variable Local Exploration
- Ensemble Kalman Variational Objectives: Nonlinear Latent Trajectory Inference with A Hybrid of Variational Inference and Ensemble Kalman Filter
- Neuro-Symbolic Representations for Video Captioning: A Case for Leveraging Inductive Biases for Vision and Language
- Wasserstein Learning of Determinantal Point Processes
- Learning Augmentation Distributions using Transformed Risk Minimization
- Improve Variational Autoencoder for Text Generationwith Discrete Latent Bottleneck
- Capsule Networks -- A Probabilistic Perspective
- Correlation-aware Unsupervised Change-point Detection via Graph Neural Networks
- DiffPrune: Neural Network Pruning with Deterministic Approximate Binary Gates and Regularization
- BiSNN: Training Spiking Neural Networks with Binary Weights via Bayesian Learning
- Support Recovery with Stochastic Gates: Theory and Application for Linear Models
- Conditioned Text Generation with Transfer for Closed-Domain Dialogue Systems
- Analyzing and Interpreting Neural Networks for NLP: A Report on the First BlackboxNLP Workshop
- Model-free Policy Learning with Reward Gradients
- Differentiable Greedy Submodular Maximization: Guarantees, Gradient Estimators, and Applications
- Network Space Search for Pareto-Efficient Spaces
- GIFnets: Differentiable GIF Encoding Framework
- Transferable Persona-Grounded Dialogues via Grounded Minimal Edits
- Coarse Grained Exponential Variational Autoencoders
- The Detection of Distributional Discrepancy for Text Generation
- Obfuscation for Privacy-preserving Syntactic Parsing
- Learning Sparsity of Representations with Discrete Latent Variables
- Generating Multi-type Temporal Sequences to Mitigate Class-imbalanced Problem
- Generating Informative Dialogue Responses with Keywords-Guided Networks
- Variational latent discrete representation for time series modelling
- Multi-Hot Compact Network Embedding
- MarioNette: Self-Supervised Sprite Learning
- Preliminary study on using vector quantization latent spaces for TTS/VC systems with consistent performance
- Learning Multi-Object Symbols for Manipulation with Attentive Deep Effect Predictors
- Multilevel Monte Carlo Variational Inference
- Differentiable Grammars for Videos
- On fine-tuning of Autoencoders for Fuzzy rule classifiers
- Topology Distillation for Recommender System
- Learning to Extend Program Graphs to Work-in-Progress Code
- Probabilistic Selective Encryption of Convolutional Neural Networks for Hierarchical Services
- Towards Faster k-Nearest-Neighbor Machine Translation
- Generalizable and Explainable Dialogue Generation via Explicit Action Learning
- Calibrate your listeners! Robust communication-based training for pragmatic speakers
- Error-Correcting Neural Sequence Prediction
- EXoN: EXplainable encoder Network
- Latte-Mix: Measuring Sentence Semantic Similarity with Latent Categorical Mixtures
- Deep Direct Likelihood Knockoffs
- Channel selection using Gumbel Softmax
- DiscoDVT: Generating Long Text with Discourse-Aware Discrete Variational Transformer
- One Network to Solve Them All: A Sequential Multi-Task Joint Learning Network Framework for MR Imaging Pipeline
- Hindsight Network Credit Assignment
- Differentiable TAN Structure Learning for Bayesian Network Classifiers
- Debiasing a First-order Heuristic for Approximate Bi-level Optimization
- Emergent symbolic language based deep medical image classification
- Modulating Scalable Gaussian Processes for Expressive Statistical Learning
- Minimizing Communication while Maximizing Performance in Multi-Agent Reinforcement Learning
- FIVES: Feature Interaction Via Edge Search for Large-Scale Tabular Data
- Disentangling Generative Factors in Natural Language with Discrete Variational Autoencoders
- Seq2Seq Mimic Games: A Signaling Perspective
- Probabilistic fine-tuning of pruning masks and PAC-Bayes self-bounded learning
- Spatially Constrained Transformer with Efficient Global Relation Modelling for Spatio-Temporal Prediction
- Bi-Directional Differentiable Input Reconstruction for Low-Resource Neural Machine Translation
- Fast Generating A Large Number of Gumbel-Max Variables
- Polyline Generative Navigable Space Segmentation for Autonomous Visual Navigation
- Approximation Based Variance Reduction for Reparameterization Gradients
- PT-Ranking: A Benchmarking Platform for Neural Learning-to-Rank
- That looks interesting! Personalizing Communication and Segmentation with Random Forest Node Embeddings
- Stein Latent Optimization for Generative Adversarial Networks
- A Comparison of Discrete Latent Variable Models for Speech Representation Learning
- Semi-supervised Learning with Contrastive Predicative Coding
- MixPoet: Diverse Poetry Generation via Learning Controllable Mixed Latent Space
- Group-disentangled Representation Learning with Weakly-Supervised Regularization
- Towards Zero-Shot Knowledge Distillation for Natural Language Processing
- Automatically Exposing Problems with Neural Dialog Models
- Variational Embeddings for Community Detection and Node Representation
- Entropy optimized semi-supervised decomposed vector-quantized variational autoencoder model based on transfer learning for multiclass text classification and generation
- ARMIN: Towards a More Efficient and Light-weight Recurrent Memory Network
- Adversarial Machine Learning in Text Analysis and Generation
- Learning sparse transformations through backpropagation
- KF-LAX: Kronecker-factored curvature estimation for control variate optimization in reinforcement learning
- Unsupervised Word Segmentation from Discrete Speech Units in Low-Resource Settings
- Adversarial Learning of Poisson Factorisation Model for Gauging Brand Sentiment in User Reviews
- Bayesian Sparsification Methods for Deep Complex-valued Networks
- Quantization Loss Re-Learning Method
- Learning Opinion Summarizers by Selecting Informative Reviews
- Self-Calibrating Indoor Localization with Crowdsourcing Fingerprints and Transfer Learning
- ABC-Di: Approximate Bayesian Computation for Discrete Data
- Lipschitz standardization for multivariate learning
- Direct-Search for a Class of Stochastic Min-Max Problems
- Effective Unsupervised Domain Adaptation with Adversarially Trained Language Models
- Learning to Expand: Reinforced Pseudo-relevance Feedback Selection for Information-seeking Conversations
- UVTomo-GAN: An adversarial learning based approach for unknown view X-ray tomographic reconstruction
- Inductive Granger Causal Modeling for Multivariate Time Series
- Learning by Examples Based on Multi-level Optimization
- An initial investigation on optimizing tandem speaker verification and countermeasure systems using reinforcement learning
- MED-TEX: Transferring and Explaining Knowledge with Less Data from Pretrained Medical Imaging Models
- Learning Natural Language Generation from Scratch
- NAST: A Non-Autoregressive Generator with Word Alignment for Unsupervised Text Style Transfer
- MSR-GAN: Multi-Segment Reconstruction via Adversarial Learning
- Generalizing Emergent Communication
- Toward Compact Parameter Representations for Architecture-Agnostic Neural Network Compression
- Risk factor identification for incident heart failure using neural network distillation and variable selection
- Dynamic Gaussian Mixture based Deep Generative Model For Robust Forecasting on Sparse Multivariate Time Series
- A Differentiable Relaxation of Graph Segmentation and Alignment for AMR Parsing
- Refining BERT Embeddings for Document Hashing via Mutual Information Maximization
- Learning source-aware representations of music in a discrete latent space
- Learning to Estimate Kernel Scale and Orientation of Defocus Blur with Asymmetric Coded Aperture
- Specializing Word Embeddings (for Parsing) by Information Bottleneck
- Searching by Generating: Flexible and Efficient One-Shot NAS with Architecture Generator
- Improvement in Machine Translation with Generative Adversarial Networks
- Learning to segment with image-level supervision
- Detect, anticipate and generate: Semi-supervised recurrent latent variable models for human activity modeling
- Semi-Relaxed Quantization with DropBits: Training Low-Bit Neural Networks via Bit-wise Regularization
- Interpretable Textual Neuron Representations for NLP
- Are You Sure You Want To Do That? Classification with Verification
- Stochastic Contrastive Learning
- PLAN-B: Predicting Likely Alternative Next Best Sequences for Action Prediction
- Asymmetrical Hierarchical Networks with Attentive Interactions for Interpretable Review-Based Recommendation
- Learning Approximately Objective Priors
- Deep Clustering of Compressed Variational Embeddings
- Learning Energy-Based Approximate Inference Networks for Structured Applications in NLP
- Hybrid Memoised Wake-Sleep: Approximate Inference at the Discrete-Continuous Interface
- NSS-VAEs: Generative Scene Decomposition for Visual Navigable Space Construction
- Differentiable Combinatorial Losses through Generalized Gradients of Linear Programs
- PatchGame: Learning to Signal Mid-level Patches in Referential Games
- Composition and decomposition of GANs
- Virtual Conditional Generative Adversarial Networks
- Regularize! Don't Mix: Multi-Agent Reinforcement Learning without Explicit Centralized Structures
- DSBERT:Unsupervised Dialogue Structure learning with BERT
- RBM-Flow and D-Flow: Invertible Flows with Discrete Energy Base Spaces
- Syntax Matters! Syntax-Controlled in Text Style Transfer
- Style Pooling: Automatic Text Style Obfuscation for Improved Classification Fairness
- Gumbel-softmax Optimization: A Simple General Framework for Combinatorial Optimization Problems on Graphs
- Learning Task Agnostic Skills with Data-driven Guidance
- Learning Modular Structures That Generalize Out-of-Distribution
- Semantic Role Labeling with Iterative Structure Refinement
- Relationships from Entity Stream
- Multi-modal dialog for browsing large visual catalogs using exploration-exploitation paradigm in a joint embedding space
- Decision Machines: Congruent Decision Trees
- SONG: Self-Organizing Neural Graphs
- An Adversarial Learning Based Approach for Unknown View Tomographic Reconstruction
- Synthetic Active Distribution System Generation via Unbalanced Graph Generative Adversarial Network
- Effect of choice of probability distribution, randomness, and search methods for alignment modeling in sequence-to-sequence text-to-speech synthesis using hard alignment
- Framing Unpacked: A Semi-Supervised Interpretable Multi-View Model of Media Frames
- Net: Augmented Parallel-Pyramid Net for Attention Guided Pose Estimation
- The Emergence of the Shape Bias Results from Communicative Efficiency
- On SkipGram Word Embedding Models with Negative Sampling: Unified Framework and Impact of Noise Distributions
- Generating the support with extreme value losses
- End-to-End Spoken Language Understanding for Generalized Voice Assistants
- Hamming Sentence Embeddings for Information Retrieval
- Lifelong Mixture of Variational Autoencoders
- NP-DRAW: A Non-Parametric Structured Latent Variable Model for Image Generation
- TaylorGAN: Neighbor-Augmented Policy Update for Sample-Efficient Natural Language Generation
- Unsupervised Learning of Depth and Deep Representation for Visual Odometry from Monocular Videos in a Metric Space
- Cause-Effect Deep Information Bottleneck For Systematically Missing Covariates
- Smaller Text Classifiers with Discriminative Cluster Embeddings
- Depthwise Discrete Representation Learning
- Medical data wrangling with sequential variational autoencoders
- A Conditional Generative Matching Model for Multi-lingual Reply Suggestion
- Learning Interpretable and Discrete Representations with Adversarial Training for Unsupervised Text Classification
- Preservation of Anomalous Subgroups On Machine Learning Transformed Data
- Minority Class Oversampling for Tabular Data with Deep Generative Models
- Counterfactual Maximum Likelihood Estimation for Training Deep Networks
- Domain-Constrained Advertising Keyword Generation
- Augment-Reinforce-Merge Policy Gradient for Binary Stochastic Policy
- Adversarial Scrubbing of Demographic Information for Text Classification
- A Bayesian Approach to Invariant Deep Neural Networks
- Towards a Better Tradeoff between Effectiveness and Efficiency in Pre-Ranking: A Learnable Feature Selection based Approach
- Sparse associative memory based on contextual code learning for disambiguating word senses
- End-to-End Learning Using Cycle Consistency for Image-to-Caption Transformations
- Plan, Attend, Generate: Character-level Neural Machine Translation with Planning in the Decoder
- Response-Anticipated Memory for On-Demand Knowledge Integration in Response Generation
- gComm: An environment for investigating generalization in Grounded Language Acquisition
- Latent Multi-Criteria Ratings for Recommendations
- Seeker: Real-Time Interactive Search
- Graph Convolutional Memory using Topological Priors
- Leveraging Hidden Structure in Self-Supervised Learning
- Illiterate DALL-E Learns to Compose
- Generalising Cost-Optimal Particle Filtering
- More Behind Your Electricity Bill: a Dual-DNN Approach to Non-Intrusive Load Monitoring