6 papers
Provable Post-Training Quantization: Theoretical Analysis of OPTQ and Qronos
Haoyu Zhang, Shihao Zhang, Ian Colbert +1
Post-training quantization (PTQ) has become a crucial tool for reducing the memory and compute costs of modern deep neural networks, including large language models (LLMs). Among P…
Qronos: Correcting the Past by Shaping the Future... in Post-Training Quantization
Shihao Zhang, Haoyu Zhang, Ian Colbert +1
We introduce Qronos -- a new state-of-the-art post-training quantization algorithm that sequentially rounds and updates neural network weights. Qronos not only explicitly corrects…
Unified Stochastic Framework for Neural Network Quantization and Pruning
Haoyu Zhang, Rayan Saab
Quantization and pruning are two essential techniques for compressing neural networks, yet they are often treated independently, with limited theoretical analysis connecting them.…
TC-RAG:Turing-Complete RAG's Case study on Medical LLM Systems
Xinke Jiang, Yue Fang, Rihong Qiu +9
In the pursuit of enhancing domain-specific Large Language Models (LLMs), Retrieval-Augmented Generation (RAG) emerges as a promising solution to mitigate issues such as hallucinat…
Dataflow-Based Optimization for Quantum Intermediate Representation Programs
Junjie Luo, Haoyu Zhang, Jianjun Zhao
This paper proposes QDFO, a dataflow-based optimization approach to Microsoft QIR. QDFO consists of two main functions: one is to preprocess the QIR code so that the LLVM optimizer…
PolyGlotFake: A Novel Multilingual and Multimodal DeepFake Dataset
Yang Hou, Haitao Fu, Chuankai Chen +3
With the rapid advancement of generative AI, multimodal deepfakes, which manipulate both audio and visual modalities, have drawn increasing public concern. Currently, deepfake dete…