6 citations · 6 across the 1 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2019
Multi-Granularity Self-Attention for Neural Machine Translation
Jie Hao, Xing Wang, Shuming Shi +2
Current state-of-the-art neural machine translation (NMT) uses a deep multi-head self-attention network with no explicit phrase information. However, prior work on statistical mach…
cs.CL2019
Towards Better Modeling Hierarchical Structure for Self-Attention with Ordered Neurons
Jie Hao, Xing Wang, Shuming Shi +2
Recent studies have shown that a hybrid of self-attention networks (SANs) and recurrent neural networks (RNNs) outperforms both individual architectures, while not much is known ab…
cs.CL2019★ 6 cited
Modeling Recurrence for Transformer
Jie Hao, Xing Wang, Baosong Yang +3
Recently, the Transformer model that is based solely on attention mechanisms, has advanced the state-of-the-art on various machine translation tasks. However, recent studies reveal…