10 citations · 20 across the 4 of their papers we have counts for
8 papers
Revisiting Hierarchical Approach for Persistent Long-Term Video Prediction
Wonkwang Lee, Whie Jung, Han Zhang +6
Learning to predict the long-term future of video frames is notoriously challenging due to inherent ambiguities in the distant future and dramatic amplifications of prediction erro…
Cost-Efficient Online Hyperparameter Optimization
Jingkang Wang, Mengye Ren, Ilija Bogunovic +2
Recent work on hyperparameters optimization (HPO) has shown the possibility of training certain hyperparameters together with regular parameters. However, these online HPO algorith…
Understanding Why Neural Networks Generalize Well Through GSNR of Parameters
Jinlong Liu, Guoqing Jiang, Yunzhi Bai +2
As deep neural networks (DNNs) achieve tremendous success across many application domains, researchers tried to explore in many aspects on why they generalize well. In this paper,…
Differentiable Product Quantization for End-to-End Embedding Compression
Ting Chen, Lala Li, Yizhou Sun
Embedding layers are commonly used to map discrete symbols into continuous embedding vectors that reflect their semantic meanings. Despite their effectiveness, the number of parame…
Doubly Sparse: Sparse Mixture of Sparse Experts for Efficient Softmax Inference
Shun Liao, Ting Chen, Tian Lin +2
Computations for the softmax function are significantly expensive when the number of output classes is large. In this paper, we present a novel softmax inference speedup method, Do…
Self-Supervised GAN to Counter Forgetting
Ting Chen, Xiaohua Zhai, Neil Houlsby
GANs involve training two networks in an adversarial game, where each network's task depends on its adversary. Recently, several works have framed GAN training as an online or cont…