127 citations · 164 across the 7 of their papers we have counts for
14 papers
Representing Long-Range Context for Graph Neural Networks with Global Attention
Zhanghao Wu, Paras Jain, Matthew A. Wright +3
Graph neural networks are powerful architectures for structured datasets. However, current methods struggle to represent long-range dependencies. Scaling the depth or width of GNNs…
Grounded Graph Decoding Improves Compositional Generalization in Question Answering
Yu Gai, Paras Jain, Wendi Zhang +3
Question answering models struggle to generalize to novel compositions of training patterns, such to longer sequences or more complex test structures. Current end-to-end models lea…
PAC Best Arm Identification Under a Deadline
Brijen Thananjeyan, Kirthevasan Kandasamy, Ion Stoica +3
We study -PAC best arm identification, where a decision-maker must identify an -optimal arm with probability at least , while minimizing the number of arm pulls (…
ActNN: Reducing Training Memory Footprint via 2-Bit Activation Compressed Training
Jianfei Chen, Lianmin Zheng, Zhewei Yao +4
The increasing size of neural network models has been critical for improvements in their accuracy, but device memory is not growing at the same rate. This creates fundamental chall…
Online Learning Demands in Max-min Fairness
Kirthevasan Kandasamy, Gur-Eyal Sela, Joseph E Gonzalez +2
We describe mechanisms for the allocation of a scarce resource among multiple users in a way that is efficient, fair, and strategy-proof, but when users do not know their resource…
Resource Allocation in Multi-armed Bandit Exploration: Overcoming Sublinear Scaling with Adaptive Parallelism
Brijen Thananjeyan, Kirthevasan Kandasamy, Ion Stoica +3
We study exploration in stochastic multi-armed bandits when we have access to a divisible resource that can be allocated in varying amounts to arm pulls. We focus in particular on…