Publications (72)
Hessian-based Analysis of Large Batch Training and Robustness to Adversaries
Zhewei Yao, Amir Gholami, Qi Lei +2
Agentic Test-Time Scaling for WebAgents
Nicholas Lee, Lutfi Eren Erdogan, Chris Joseph John +4
Applications and Techniques for Fast Machine Learning in Science
Allison McCarn Deiana, Nhan Tran, Joshua Agar +84
HAWQ-V2: Hessian Aware trace-Weighted Quantization of Neural Networks
Zhen Dong, Zhewei Yao, Yaohui Cai +4
Learned Token Pruning for Transformers
Sehoon Kim, Sheng Shen, David Thorsley +4
ANODEV2: A Coupled Neural ODE Evolution Framework
Tianjun Zhang, Zhewei Yao, Amir Gholami +4
Adaptive Self-supervision Algorithms for Physics-informed Neural Networks
Shashank Subramanian, Robert M. Kirby, Michael W. Mahoney +1
Parameter Re-Initialization through Cyclical Batch Size Schedules
Norman Mu, Zhewei Yao, Amir Gholami +2
FFT, FMM, or Multigrid? A comparative Study of State-Of-the-Art Poisson Solvers for Uniform and Nonuniform Grids in the Unit Cube
Amir Gholami, Dhairya Malhotra, Hari Sundar +1
Integer-only Zero-shot Quantization for Efficient Speech Recognition
Sehoon Kim, Amir Gholami, Zhewei Yao +7
Full Stack Optimization of Transformer Inference: a Survey
Sehoon Kim, Coleman Hooper, Thanakul Wattanawong +9
On Neural Scaling Laws for Weather Emulation through Continual Training
Shashank Subramanian, Alexander Kiefer, Arnur Nigmetov +3
Speculative Decoding with Big Little Decoder
Sehoon Kim, Karttikeya Mangalam, Suhong Moon +4
Characterizing Prompt Compression Methods for Long Context Inference
Siddharth Jha, Lutfi Eren Erdogan, Sehoon Kim +2
Towards Foundation Models for Scientific Machine Learning: Characterizing Scaling and Transfer Behavior
Shashank Subramanian, Peter Harrington, Kurt Keutzer +4
An LLM Compiler for Parallel Function Calling
Sehoon Kim, Suhong Moon, Ryan Tabrizi +4
CLAIRE: A distributed-memory solver for constrained large deformation diffeomorphic image registration
Andreas Mang, Amir Gholami, Christos Davatzikos +1
Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
Lutfi Eren Erdogan, Nicholas Lee, Sehoon Kim +5
SqueezeNext: Hardware-Aware Neural Network Design
Amir Gholami, Kiseok Kwon, Bichen Wu +5
Speculative Interaction Agents: Building Real-Time Agents with Asynchronous I/O and Speculative Tool Calling
Coleman Hooper, Minwoo Kang, Suhong Moon +7
Multipole Attention for Efficient Long Context Reasoning
Coleman Hooper, Sebastian Zhao, Luca Manolache +5
Squeezeformer: An Efficient Transformer for Automatic Speech Recognition
Sehoon Kim, Amir Gholami, Albert Shaw +5
LoSA: Locality Aware Sparse Attention for Block-Wise Diffusion Language Models
Haocheng Xi, Harman Singh, Yuezhou Hu +9
XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization
Aditya Tomar, Coleman Hooper, Minjae Lee +7
Simulation of glioblastoma growth using a 3D multispecies tumor model with mass effect
Shashank Subramanian, Amir Gholami, George Biros
A Fast Post-Training Pruning Framework for Transformers
Woosuk Kwon, Sehoon Kim, Michael W. Mahoney +3
Distributed-memory large deformation diffeomorphic 3D image registration
Andreas Mang, Amir Gholami, George Biros
HAWQ: Hessian AWare Quantization of Neural Networks with Mixed-Precision
Zhen Dong, Zhewei Yao, Amir Gholami +2
CDLM: Consistency Diffusion Language Models For Faster Sampling
Minseo Kim, Chenfeng Xu, Coleman Hooper +5
Squeezed Attention: Accelerating Long Context Length LLM Inference
Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh +6
AI and Memory Wall
Amir Gholami, Zhewei Yao, Sehoon Kim +3
Checkmate: Breaking the Memory Wall with Optimal Tensor Rematerialization
Paras Jain, Ajay Jain, Aniruddha Nrusimha +5
Arbitrage: Efficient Reasoning via Advantage-Aware Speculation
Monishwaran Maheswaran, Rishabh Tiwari, Yuezhou Hu +8
Characterizing possible failure modes in physics-informed neural networks
Aditi S. Krishnapriyan, Amir Gholami, Shandian Zhe +2
PowerNorm: Rethinking Batch Normalization in Transformers
Sheng Shen, Zhewei Yao, Amir Gholami +2
ADAHESSIAN: An Adaptive Second Order Optimizer for Machine Learning
Zhewei Yao, Amir Gholami, Sheng Shen +3
Identifying the Best Machine Learning Algorithms for Brain Tumor Segmentation, Progression Assessment, and Overall Survival Prediction in the BRATS Challenge
Spyridon Bakas, Mauricio Reyes, Andras Jakab +421
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
Rishabh Tiwari, Haocheng Xi, Aditya Tomar +7
Trust Region Based Adversarial Attack on Neural Networks
Zhewei Yao, Amir Gholami, Peng Xu +2
On the Computational Inefficiency of Large Batch Sizes for Stochastic Gradient Descent
Noah Golmant, Nikita Vemuri, Zhewei Yao +5
SqueezeLLM: Dense-and-Sparse Quantization
Sehoon Kim, Coleman Hooper, Amir Gholami +5
Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
Minseo Kim, Coleman Hooper, Aditya Tomar +5
AccFFT: A library for distributed-memory FFT on CPU and GPU architectures
Amir Gholami, Judith Hill, Dhairya Malhotra +1
An inverse problem formulation for parameter estimation of a reaction diffusion model of low grade gliomas
Amir Gholami, Andreas Mang, George Biros
Coupling Brain-Tumor Biophysical Models and Diffeomorphic Image Registration
Klaudius Scheufele, Andreas Mang, Amir Gholami +3
A Survey of Quantization Methods for Efficient Neural Network Inference
Amir Gholami, Sehoon Kim, Zhen Dong +3
Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT
Sheng Shen, Zhen Dong, Jiayu Ye +5
EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts
Minseo Kim, Minjae Lee, Seunghyuk Oh +7
Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models
Rishabh Tiwari, Aditya Tomar, Udbhav Bamba +5
TinyAgent: Function Calling at the Edge
Lutfi Eren Erdogan, Nicholas Lee, Siddharth Jha +7
Inefficiency of K-FAC for Large Batch Size Training
Linjian Ma, Gabe Montague, Jiayu Ye +4
Boundary thickness and robustness in learning models
Yaoqing Yang, Rajiv Khanna, Yaodong Yu +5
End-to-end codesign of Hessian-aware quantized neural networks for FPGAs and ASICs
Javier Campos, Zhen Dong, Javier Duarte +4
Integrated Model, Batch and Domain Parallelism in Training Neural Networks
Amir Gholami, Ariful Azad, Peter Jin +2
Hessian-Aware Pruning and Optimal Neural Implant
Shixing Yu, Zhewei Yao, Amir Gholami +4
Large batch size training of neural networks with adversarial training and second-order information
Zhewei Yao, Amir Gholami, Daiyaan Arfeen +4
A Novel Domain Adaptation Framework for Medical Image Segmentation
Amir Gholami, Shashank Subramanian, Varun Shenoy +6
Residual Context Diffusion Language Models
Yuezhou Hu, Harman Singh, Monishwaran Maheswaran +10
ETS: Efficient Tree Search for Inference-Time Scaling
Coleman Hooper, Sehoon Kim, Suhong Moon +7
HAWQV3: Dyadic Neural Network Quantization
Zhewei Yao, Zhen Dong, Zhangcheng Zheng +8
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh +4
LLM2LLM: Boosting LLMs with Novel Iterative Data Enhancement
Nicholas Lee, Thanakul Wattanawong, Sehoon Kim +6
ZeroQ: A Novel Zero Shot Quantization Framework
Yaohui Cai, Zhewei Yao, Zhen Dong +3
PDE-constrained optimization in medical image analysis
Andreas Mang, Amir Gholami, Christos Davatzikos +1
Co-Design of Deep Neural Nets and Neural Net Accelerators for Embedded Vision Applications
Kiseok Kwon, Alon Amid, Amir Gholami +3
I-BERT: Integer-only BERT Quantization
Sehoon Kim, Amir Gholami, Zhewei Yao +2
SciML Agents: Write the Solver, Not the Solution
Saarth Gaonkar, Xiang Zheng, Haocheng Xi +5
Reliable edge machine learning hardware for scientific applications
Tommaso Baldi, Javier Campos, Ben Hawks +15
ANODE: Unconditionally Accurate Memory-Efficient Gradients for Neural ODEs
Amir Gholami, Kurt Keutzer, George Biros
SPEED: Speculative Pipelined Execution for Efficient Decoding
Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh +4
Efficient and Scalable Estimation of Tool Representations in Vector Space
Suhong Moon, Siddharth Jha, Lutfi Eren Erdogan +4
PyHessian: Neural Networks Through the Lens of the Hessian
Zhewei Yao, Amir Gholami, Kurt Keutzer +1