79 citations · 81 across the 4 of their papers we have counts for
13 papers · 1 filter
Efficient and Scalable Estimation of Tool Representations in Vector Space
Suhong Moon, Siddharth Jha, Lutfi Eren Erdogan +4
Recent advancements in function calling and tool use have significantly enhanced the capabilities of large language models (LLMs) by enabling them to interact with external informa…
AI and Memory Wall
Amir Gholami, Zhewei Yao, Sehoon Kim +3
The availability of unprecedented unsupervised training data, along with neural scaling laws, has resulted in an unprecedented surge in model size and compute requirements for serv…
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh +4
LLMs are seeing growing use for applications which require large context windows, and with these large context windows KV cache activations surface as the dominant contributor to m…
Boundary thickness and robustness in learning models
Yaoqing Yang, Rajiv Khanna, Yaodong Yu +5
Robustness of machine learning models to various adversarial and non-adversarial corruptions continues to be of interest. In this paper, we introduce the notion of the boundary thi…
PyHessian: Neural Networks Through the Lens of the Hessian
Zhewei Yao, Amir Gholami, Kurt Keutzer +1
We present PYHESSIAN, a new scalable framework that enables fast computation of Hessian (i.e., second-order derivative) information for deep neural networks. PYHESSIAN enables fast…
Checkmate: Breaking the Memory Wall with Optimal Tensor Rematerialization
Paras Jain, Ajay Jain, Aniruddha Nrusimha +5
We formalize the problem of trading-off DNN training time and memory requirements as the tensor rematerialization optimization problem, a generalization of prior checkpointing stra…