47 citations · 52 across the 11 of their papers we have counts for
4 papers · 1 filter
Towards Efficient Verification of Quantized Neural Networks
Pei Huang, Haoze Wu, Yuting Yang +4
Quantization replaces floating point arithmetic with integer arithmetic in deep neural network models, providing more efficient on-device inference with less power and memory. In t…
Convex Bounds on the Softmax Function with Applications to Robustness Verification
Dennis Wei, Haoze Wu, Min Wu +3
The softmax function is a ubiquitous component at the output of neural networks and increasingly in intermediate layers as well. This paper provides convex lower bounds and concave…
FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Ying Sheng, Lianmin Zheng, Binhang Yuan +11
The high computational and memory requirements of large language model (LLM) inference make it feasible only with multiple high-end accelerators. Motivated by the emerging demand f…
On Optimizing Back-Substitution Methods for Neural Network Verification
Tom Zelazny, Haoze Wu, Clark Barrett +1
With the increasing application of deep learning in mission-critical systems, there is a growing need to obtain formal guarantees about the behaviors of neural networks. Indeed, ma…