42 citations · 81 across the 10 of their papers we have counts for
15 papers
Surrogate Gap Minimization Improves Sharpness-Aware Training
Juntang Zhuang, Boqing Gong, Liangzhe Yuan +6
The recently proposed Sharpness-Aware Minimization (SAM) improves generalization by minimizing a \textit{perturbed loss} defined as the maximum loss within a neighborhood in the pa…
MALI: A memory efficient and reverse accurate integrator for Neural ODEs
Juntang Zhuang, Nicha C. Dvornek, Sekhar Tatikonda +1
Neural ordinary differential equations (Neural ODEs) are a new family of deep-learning models with continuous depth. However, the numerical estimation of the gradient in the contin…
Multiple-shooting adjoint method for whole-brain dynamic causal modeling
Juntang Zhuang, Nicha Dvornek, Sekhar Tatikonda +3
Dynamic causal modeling (DCM) is a Bayesian framework to infer directed connections between compartments, and has been used to describe the interactions between underlying neural p…
AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed Gradients
Juntang Zhuang, Tommy Tang, Yifan Ding +4
Most popular optimizers for deep learning can be broadly categorized as adaptive methods (e.g. Adam) and accelerated schemes (e.g. stochastic gradient descent (SGD) with momentum).…
Adaptive Checkpoint Adjoint Method for Gradient Estimation in Neural ODE
Juntang Zhuang, Nicha Dvornek, Xiaoxiao Li +3
Neural ordinary differential equations (NODEs) have recently attracted increasing attention; however, their empirical performance on benchmark tasks (e.g. image classification) are…
Sparse Regression Codes
Ramji Venkataramanan, Sekhar Tatikonda, Andrew Barron
Developing computationally-efficient codes that approach the Shannon-theoretic limits for communication and compression has long been one of the major goals of information and codi…