Order Matters: Sequence to sequence for sets
arXiv:1511.06391
Abstract
Sequences have become first class citizens in supervised learning thanks to the resurgence of recurrent neural networks. Many complex tasks that require mapping from or to a sequence of observations can now be formulated with the sequence-to-sequence (seq2seq) framework which employs the chain rule to efficiently represent the joint probability of sequences. In many cases, however, variable sized inputs and/or outputs might not be naturally expressed as sequences. For instance, it is not clear how to input a set of numbers into a model where the task is to sort them; similarly, we do not know how to organize outputs when they correspond to random variables and the task is to model their unknown joint probability. In this paper, we first show using various examples that the order in which we organize input and/or output data matters significantly when learning an underlying model. We then discuss an extension of the seq2seq framework that goes beyond sequences and handles input sets in a principled way. In addition, we propose a loss which, by searching over possible orders during training, deals with the lack of structure of output sets. We show empirical evidence of our claims regarding ordering, and on the modifications to the seq2seq framework on benchmark language modeling and parsing tasks, as well as two artificial tasks -- sorting numbers and estimating the joint probability of unknown graphical models.
Accepted as a conference paper at ICLR 2015
References in corpus (4)
Cited by in corpus (94)
- PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation
- PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space
- Graph Networks as a Universal Machine Learning Framework for Molecules and Crystals
- Graph Neural Networks: A Review of Methods and Applications
- Matching Networks for One Shot Learning
- Hierarchical Graph Representation Learning with Differentiable Pooling
- Density estimation using Real NVP
- Learning Deep Generative Models of Graphs
- Few-Shot Learning with Graph Neural Networks
- Adaptive Computation Time for Recurrent Neural Networks
- Principal Neighbourhood Aggregation for Graph Nets
- Robust Attentional Aggregation of Deep Feature Sets for Multi-view 3D Reconstruction
- Hierarchical Graph Pooling with Structure Learning
- Understanding Pooling in Graph Neural Networks
- Dual Attention Matching Network for Context-Aware Feature Sequence based Person Re-Identification
- Compositional Generalization in Semantic Parsing: Pre-training vs. Specialized Architectures
- HighRes-net: Recursive Fusion for Multi-Frame Super-Resolution of Satellite Imagery
- Learning the Multiple Traveling Salesmen Problem with Permutation Invariant Pooling Networks
- On Learning Sets of Symmetric Elements
- MoleculeNet: A Benchmark for Molecular Machine Learning
- All SMILES Variational Autoencoder
- Table-to-text Generation by Structure-aware Seq2seq Learning
- Learning Permutations with Sinkhorn Policy Gradient
- Self-Supervised Deep Learning on Point Clouds by Reconstructing Space
- Neural Language Generation: Formulation, Methods, and Evaluation
- What to talk about and how? Selective Generation using LSTMs with Coarse-to-Fine Alignment
- Unsupervised Attributed Multiplex Network Embedding
- Sememe Prediction: Learning Semantic Knowledge from Unstructured Textual Wiki Descriptions
- Searching for Effective Neural Extractive Summarization: What Works and What's Next
- Exploration on Generating Traditional Chinese Medicine Prescription from Symptoms with an End-to-End method
- Strong Generalization and Efficiency in Neural Programs
- Pointer Graph Networks
- The Emergence of Compositional Languages for Numeric Concepts Through Iterated Learning in Neural Agents
- One-Shot Relational Learning for Knowledge Graphs
- Utilizing Edge Features in Graph Neural Networks via Variational Information Maximization
- Set2Graph: Learning Graphs From Sets
- Exchangeable deep neural networks for set-to-set matching and learning
- Learning a Simple and Effective Model for Multi-turn Response Generation with Auxiliary Tasks
- Graph Deconvolutional Generation
- A Novel Genetic Algorithm with Hierarchical Evaluation Strategy for Hyperparameter Optimisation of Graph Neural Networks
- Discriminative structural graph classification
- Iterative Graph Self-Distillation
- Exchangeable Neural ODE for Set Modeling
- Fast Adaptation in Generative Models with Generative Matching Networks
- Learning Graph-Level Representations with Recurrent Neural Networks
- Graph Convolutional Networks with EigenPooling
- An Entity-Driven Framework for Abstractive Summarization
- Dependency-aware Attention Control for Unconstrained Face Recognition with Image Sets
- Self-supervised edge features for improved Graph Neural Network training
- A Deep Reinforced Sequence-to-Set Model for Multi-Label Text Classification
- Neural Message Passing for Multi-Label Classification
- EvoNet: A Neural Network for Predicting the Evolution of Dynamic Graphs
- Towards Solving Text-based Games by Producing Adaptive Action Spaces
- EventNet: Asynchronous Recursive Event Processing
- IPC-Net: 3D point-cloud segmentation using deep inter-point convolutional layers
- Enhancing Label Correlation Feedback in Multi-Label Text Classification via Multi-Task Learning
- Reconstruction for Powerful Graph Representations
- Maximum Entropy Weighted Independent Set Pooling for Graph Neural Networks
- Better Set Representations For Relational Reasoning
- CommPOOL: An Interpretable Graph Pooling Framework for Hierarchical Graph Representation Learning
- Accurate Prediction of Free Solvation Energy of Organic Molecules via Graph Attention Network and Message Passing Neural Network from Pairwise Atomistic Interactions
- Emergence of Numeric Concepts in Multi-Agent Autonomous Communication
- Learning with Sets in Multiple Instance Regression Applied to Remote Sensing
- Learning Hierarchical Review Graph Representations for Recommendation
- Learning to generate classifiers
- Chained Predictions Using Convolutional Neural Networks
- Few-Shot Object Recognition from Machine-Labeled Web Images
- Exact-K Recommendation via Maximal Clique Optimization
- Neuralizing Efficient Higher-order Belief Propagation
- Predicting Kovats Retention Indices Using Graph Neural Networks
- CGCL: Collaborative Graph Contrastive Learning without Handcrafted Graph Data Augmentations
- Compositional Embeddings for Multi-Label One-Shot Learning
- Learning Set-equivariant Functions with SWARM Mappings
- Finite Group Equivariant Neural Networks for Games
- Learning Sentence Embeddings for Coherence Modelling and Beyond
- SGVAE: Sequential Graph Variational Autoencoder
- Seq-SetNet: Exploring Sequence Sets for Inferring Structures
- Diversified Multiscale Graph Learning with Graph Self-Correction
- RaWaNet: Enriching Graph Neural Network Input via Random Walks on Graphs
- Kernel Mean Embedding of Instance-wise Predictions in Multiple Instance Regression
- ProDyn0: Inferring calponin homology domain stretching behavior using graph neural networks
- Sequential Graph Dependency Parser
- Rethinking PointNet Embedding for Faster and Compact Model
- Knowledge as Invariance -- History and Perspectives of Knowledge-augmented Machine Learning
- Understood in Translation, Transformers for Domain Understanding
- Learning Indoor Layouts from Simple Point-Clouds
- Content-Based Table Retrieval for Web Queries
- Here's My Point: Joint Pointer Architecture for Argument Mining
- The CAT SET on the MAT: Cross Attention for Set Matching in Bipartite Hypergraphs
- Differentiable Representations For Multihop Inference Rules
- Using holistic event information in the trigger
- One-Shot Learning for Language Modelling
- OrderNet: Ordering by Example
- Genetic Constrained Graph Variational Autoencoder for COVID-19 Drug Discovery