NewEvery arXiv paper, its researchers & institutions — mapped.
papers

Publications (72)

cs.CV2018

Hessian-based Analysis of Large Batch Training and Robustness to Adversaries

Zhewei Yao, Amir Gholami, Qi Lei +2

cs.AI2026

Agentic Test-Time Scaling for WebAgents

Nicholas Lee, Lutfi Eren Erdogan, Chris Joseph John +4

cs.LG2021

Applications and Techniques for Fast Machine Learning in Science

Allison McCarn Deiana, Nhan Tran, Joshua Agar +84

cs.CV2019

HAWQ-V2: Hessian Aware trace-Weighted Quantization of Neural Networks

Zhen Dong, Zhewei Yao, Yaohui Cai +4

cs.CL2022

Learned Token Pruning for Transformers

Sehoon Kim, Sheng Shen, David Thorsley +4

cs.LG2019

ANODEV2: A Coupled Neural ODE Evolution Framework

Tianjun Zhang, Zhewei Yao, Amir Gholami +4

cs.LG2022

Adaptive Self-supervision Algorithms for Physics-informed Neural Networks

Shashank Subramanian, Robert M. Kirby, Michael W. Mahoney +1

cs.LG2018

Parameter Re-Initialization through Cyclical Batch Size Schedules

Norman Mu, Zhewei Yao, Amir Gholami +2

math.NA2016

FFT, FMM, or Multigrid? A comparative Study of State-Of-the-Art Poisson Solvers for Uniform and Nonuniform Grids in the Unit Cube

Amir Gholami, Dhairya Malhotra, Hari Sundar +1

eess.AS2022

Integer-only Zero-shot Quantization for Efficient Speech Recognition

Sehoon Kim, Amir Gholami, Zhewei Yao +7

cs.CL2023

Full Stack Optimization of Transformer Inference: a Survey

Sehoon Kim, Coleman Hooper, Thanakul Wattanawong +9

cs.LG2026

On Neural Scaling Laws for Weather Emulation through Continual Training

Shashank Subramanian, Alexander Kiefer, Arnur Nigmetov +3

cs.CL2023

Speculative Decoding with Big Little Decoder

Sehoon Kim, Karttikeya Mangalam, Suhong Moon +4

cs.CL2024

Characterizing Prompt Compression Methods for Long Context Inference

Siddharth Jha, Lutfi Eren Erdogan, Sehoon Kim +2

cs.LG2023

Towards Foundation Models for Scientific Machine Learning: Characterizing Scaling and Transfer Behavior

Shashank Subramanian, Peter Harrington, Kurt Keutzer +4

cs.CL2024

An LLM Compiler for Parallel Function Calling

Sehoon Kim, Suhong Moon, Ryan Tabrizi +4

math.OC2019

CLAIRE: A distributed-memory solver for constrained large deformation diffeomorphic image registration

Andreas Mang, Amir Gholami, Christos Davatzikos +1

cs.CL2025

Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks

Lutfi Eren Erdogan, Nicholas Lee, Sehoon Kim +5

cs.NE2018

SqueezeNext: Hardware-Aware Neural Network Design

Amir Gholami, Kiseok Kwon, Bichen Wu +5

cs.LG2026

Speculative Interaction Agents: Building Real-Time Agents with Asynchronous I/O and Speculative Tool Calling

Coleman Hooper, Minwoo Kang, Suhong Moon +7

cs.CL2025

Multipole Attention for Efficient Long Context Reasoning

Coleman Hooper, Sebastian Zhao, Luca Manolache +5

eess.AS2022

Squeezeformer: An Efficient Transformer for Automatic Speech Recognition

Sehoon Kim, Amir Gholami, Albert Shaw +5

cs.CL2026

LoSA: Locality Aware Sparse Attention for Block-Wise Diffusion Language Models

Haocheng Xi, Harman Singh, Yuezhou Hu +9

cs.LG2025

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization

Aditya Tomar, Coleman Hooper, Minjae Lee +7

physics.med-ph2019

Simulation of glioblastoma growth using a 3D multispecies tumor model with mass effect

Shashank Subramanian, Amir Gholami, George Biros

cs.CL2022

A Fast Post-Training Pruning Framework for Transformers

Woosuk Kwon, Sehoon Kim, Michael W. Mahoney +3

cs.DC2016

Distributed-memory large deformation diffeomorphic 3D image registration

Andreas Mang, Amir Gholami, George Biros

cs.CV2019

HAWQ: Hessian AWare Quantization of Neural Networks with Mixed-Precision

Zhen Dong, Zhewei Yao, Amir Gholami +2

cs.LG2026

CDLM: Consistency Diffusion Language Models For Faster Sampling

Minseo Kim, Chenfeng Xu, Coleman Hooper +5

cs.CL2025

Squeezed Attention: Accelerating Long Context Length LLM Inference

Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh +6

cs.LG2024

AI and Memory Wall

Amir Gholami, Zhewei Yao, Sehoon Kim +3

cs.LG2020

Checkmate: Breaking the Memory Wall with Optimal Tensor Rematerialization

Paras Jain, Ajay Jain, Aniruddha Nrusimha +5

cs.CL2025

Arbitrage: Efficient Reasoning via Advantage-Aware Speculation

Monishwaran Maheswaran, Rishabh Tiwari, Yuezhou Hu +8

cs.LG2021

Characterizing possible failure modes in physics-informed neural networks

Aditi S. Krishnapriyan, Amir Gholami, Shandian Zhe +2

cs.CL2020

PowerNorm: Rethinking Batch Normalization in Transformers

Sheng Shen, Zhewei Yao, Amir Gholami +2

cs.LG2021

ADAHESSIAN: An Adaptive Second Order Optimizer for Machine Learning

Zhewei Yao, Amir Gholami, Sheng Shen +3

cs.CV2019

Identifying the Best Machine Learning Algorithms for Brain Tumor Segmentation, Progression Assessment, and Overall Survival Prediction in the BRATS Challenge

Spyridon Bakas, Mauricio Reyes, Andras Jakab +421

cs.LG2025

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache

Rishabh Tiwari, Haocheng Xi, Aditya Tomar +7

cs.LG2018

Trust Region Based Adversarial Attack on Neural Networks

Zhewei Yao, Amir Gholami, Peng Xu +2

cs.LG2018

On the Computational Inefficiency of Large Batch Sizes for Stochastic Gradient Descent

Noah Golmant, Nikita Vemuri, Zhewei Yao +5

cs.CL2024

SqueezeLLM: Dense-and-Sparse Quantization

Sehoon Kim, Coleman Hooper, Amir Gholami +5

cs.LG2025

Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models

Minseo Kim, Coleman Hooper, Aditya Tomar +5

cs.DC2016

AccFFT: A library for distributed-memory FFT on CPU and GPU architectures

Amir Gholami, Judith Hill, Dhairya Malhotra +1

math.NA2015

An inverse problem formulation for parameter estimation of a reaction diffusion model of low grade gliomas

Amir Gholami, Andreas Mang, George Biros

physics.med-ph2018

Coupling Brain-Tumor Biophysical Models and Diffeomorphic Image Registration

Klaudius Scheufele, Andreas Mang, Amir Gholami +3

cs.CV2021

A Survey of Quantization Methods for Efficient Neural Network Inference

Amir Gholami, Sehoon Kim, Zhen Dong +3

cs.CL2019

Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT

Sheng Shen, Zhen Dong, Jiayu Ye +5

cs.LG2026

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts

Minseo Kim, Minjae Lee, Seunghyuk Oh +7

cs.LG2026

Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models

Rishabh Tiwari, Aditya Tomar, Udbhav Bamba +5

cs.CL2024

TinyAgent: Function Calling at the Edge

Lutfi Eren Erdogan, Nicholas Lee, Siddharth Jha +7

cs.LG2019

Inefficiency of K-FAC for Large Batch Size Training

Linjian Ma, Gabe Montague, Jiayu Ye +4

cs.LG2021

Boundary thickness and robustness in learning models

Yaoqing Yang, Rajiv Khanna, Yaodong Yu +5

cs.LG2023

End-to-end codesign of Hessian-aware quantized neural networks for FPGAs and ASICs

Javier Campos, Zhen Dong, Javier Duarte +4

cs.LG2018

Integrated Model, Batch and Domain Parallelism in Training Neural Networks

Amir Gholami, Ariful Azad, Peter Jin +2

cs.CV2021

Hessian-Aware Pruning and Optimal Neural Implant

Shixing Yu, Zhewei Yao, Amir Gholami +4

cs.LG2020

Large batch size training of neural networks with adversarial training and second-order information

Zhewei Yao, Amir Gholami, Daiyaan Arfeen +4

cs.CV2018

A Novel Domain Adaptation Framework for Medical Image Segmentation

Amir Gholami, Shashank Subramanian, Varun Shenoy +6

cs.CL2026

Residual Context Diffusion Language Models

Yuezhou Hu, Harman Singh, Monishwaran Maheswaran +10

cs.LG2025

ETS: Efficient Tree Search for Inference-Time Scaling

Coleman Hooper, Sehoon Kim, Suhong Moon +7

cs.CV2021

HAWQV3: Dyadic Neural Network Quantization

Zhewei Yao, Zhen Dong, Zhangcheng Zheng +8

cs.LG2025

KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh +4

cs.CL2024

LLM2LLM: Boosting LLMs with Novel Iterative Data Enhancement

Nicholas Lee, Thanakul Wattanawong, Sehoon Kim +6

cs.CV2020

ZeroQ: A Novel Zero Shot Quantization Framework

Yaohui Cai, Zhewei Yao, Zhen Dong +3

math.OC2018

PDE-constrained optimization in medical image analysis

Andreas Mang, Amir Gholami, Christos Davatzikos +1

cs.DC2018

Co-Design of Deep Neural Nets and Neural Net Accelerators for Embedded Vision Applications

Kiseok Kwon, Alon Amid, Amir Gholami +3

cs.CL2021

I-BERT: Integer-only BERT Quantization

Sehoon Kim, Amir Gholami, Zhewei Yao +2

cs.LG2025

SciML Agents: Write the Solver, Not the Solution

Saarth Gaonkar, Xiang Zheng, Haocheng Xi +5

cs.LG2024

Reliable edge machine learning hardware for scientific applications

Tommaso Baldi, Javier Campos, Ben Hawks +15

cs.LG2019

ANODE: Unconditionally Accurate Memory-Efficient Gradients for Neural ODEs

Amir Gholami, Kurt Keutzer, George Biros

cs.CL2024

SPEED: Speculative Pipelined Execution for Efficient Decoding

Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh +4

cs.LG2024

Efficient and Scalable Estimation of Tool Representations in Vector Space

Suhong Moon, Siddharth Jha, Lutfi Eren Erdogan +4

cs.LG2020

PyHessian: Neural Networks Through the Lens of the Hessian

Zhewei Yao, Amir Gholami, Kurt Keutzer +1