Self-attention with Functional Time Representation Learning
arXiv:1911.12864
Abstract
Sequential modelling with self-attention has achieved cutting edge performances in natural language processing. With advantages in model flexibility, computation complexity and interpretability, self-attention is gradually becoming a key component in event sequence models. However, like most other sequence models, self-attention does not account for the time span between events and thus captures sequential signals rather than temporal patterns. Without relying on recurrent network structures, self-attention recognizes event orderings via positional encoding. To bridge the gap between modelling time-independent and time-dependent event sequence, we introduce a functional feature map that embeds time span into high-dimensional spaces. By constructing the associated translation-invariant time kernel function, we reveal the functional forms of the feature map under classic functional function analysis results, namely Bochner's Theorem and Mercer's Theorem. We propose several models to learn the functional time representation and the interactions with event representation. These methods are evaluated on real-world datasets under various continuous-time event sequence prediction tasks. The experiments reveal that the proposed methods compare favorably to baseline models while also capturing useful time-event interactions.
Cited by in corpus (12)
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains
- Foundations and modelling of dynamic networks using Dynamic Graph Neural Networks: A survey
- Inductive Representation Learning in Temporal Networks via Causal Anonymous Walks
- Multi-Time Attention Networks for Irregularly Sampled Time Series
- SAPE: Spatially-Adaptive Progressive Encoding for Neural Optimization
- A Survey on Principles, Models and Methods for Learning from Irregularly Sampled Time Series
- Continuous-Time Sequential Recommendation with Temporal Graph Collaborative Transformer
- VigDet: Knowledge Informed Neural Temporal Point Process for Coordination Detection on Social Media
- Learning Self-Modulating Attention in Continuous Time Space with Applications to Sequential Recommendation
- ConTIG: Continuous Representation Learning on Temporal Interaction Graphs
- Identifying Coordinated Accounts on Social Media through Hidden Influence and Group Behaviours
- A Temporal Kernel Approach for Deep Learning with Continuous-time Information