activity
20182022
most citedLuna: Linear Unified Nested Attention

49 citations · 97 across the 9 of their papers we have counts for

collaborators

15 papers

cs.CL202217 cited

In-context Examples Selection for Machine Translation

Sweta Agrawal, Chunting Zhou, Mike Lewis +2

Large-scale generative models show an impressive ability to perform a wide range of Natural Language Processing (NLP) tasks using in-context learning, where a few examples are used…

cs.HC2022

Incorporation of Human Knowledge into Data Embeddings to Improve Pattern Significance and Interpretability

Jie Li, Chun-qi Zhou

Embedding is a common technique for analyzing multi-dimensional data. However, the embedding projection cannot always form significant and interpretable visual structures that fore…

cs.CL2021

Distributionally Robust Multilingual Machine Translation

Chunting Zhou, Daniel Levy, Xian Li +2

Multilingual neural machine translation (MNMT) learns to translate multiple language pairs with a single model, potentially improving both the accuracy and the memory-efficiency of…

cs.LG20213 cited

Examining and Combating Spurious Features under Distribution Shift

Chunting Zhou, Xuezhe Ma, Paul Michel +1

A central goal of machine learning is to learn robust representations that capture the causal relationship between inputs features and output labels. However, minimizing empirical…

cs.LG202149 cited

Luna: Linear Unified Nested Attention

Xuezhe Ma, Xiang Kong, Sinong Wang +4

The quadratic computational and memory complexities of the Transformer's attention mechanism have limited its scalability for modeling long sequences. In this paper, we propose Lun…

cs.LG2021

Learning Structures for Deep Neural Networks

Jinhui Yuan, Fei Pan, Chunting Zhou +2

In this paper, we focus on the unsupervised setting for structure learning of deep neural networks and propose to adopt the efficient coding principle, rooted in information theory…