activity
20242026
collaborators

6 papers

cs.LG2026

Provable Post-Training Quantization: Theoretical Analysis of OPTQ and Qronos

Haoyu Zhang, Shihao Zhang, Ian Colbert +1

Post-training quantization (PTQ) has become a crucial tool for reducing the memory and compute costs of modern deep neural networks, including large language models (LLMs). Among P…

cs.LG2026

Qronos: Correcting the Past by Shaping the Future... in Post-Training Quantization

Shihao Zhang, Haoyu Zhang, Ian Colbert +1

We introduce Qronos -- a new state-of-the-art post-training quantization algorithm that sequentially rounds and updates neural network weights. Qronos not only explicitly corrects…

cs.LG2025

Unified Stochastic Framework for Neural Network Quantization and Pruning

Haoyu Zhang, Rayan Saab

Quantization and pruning are two essential techniques for compressing neural networks, yet they are often treated independently, with limited theoretical analysis connecting them.…

cs.IR2024

TC-RAG:Turing-Complete RAG's Case study on Medical LLM Systems

Xinke Jiang, Yue Fang, Rihong Qiu +9

In the pursuit of enhancing domain-specific Large Language Models (LLMs), Retrieval-Augmented Generation (RAG) emerges as a promising solution to mitigate issues such as hallucinat…

cs.PL2024

Dataflow-Based Optimization for Quantum Intermediate Representation Programs

Junjie Luo, Haoyu Zhang, Jianjun Zhao

This paper proposes QDFO, a dataflow-based optimization approach to Microsoft QIR. QDFO consists of two main functions: one is to preprocess the QIR code so that the LLVM optimizer…

cs.SD2024

PolyGlotFake: A Novel Multilingual and Multimodal DeepFake Dataset

Yang Hou, Haitao Fu, Chuankai Chen +3

With the rapid advancement of generative AI, multimodal deepfakes, which manipulate both audio and visual modalities, have drawn increasing public concern. Currently, deepfake dete…