4 papers
Contextual Utility of Quantization Moves in Extreme Low-Bit LLMs
Wenxuan Xiao, Xu Cao
Post-training quantizers select finite code changes using reconstruction proxies or local loss approximations, but the utility of a quantization move depends on the state through w…
When Does Low-Bit Quantization Preserve the Decisions of Vector Search?
Wenxuan Xiao, Xu Cao
Low-bit quantization can achieve high recall on some vector representations and fail sharply on others, while average distortion and global rank correlation do not explain the diff…
Covariance Structure and Coordinate Heterogeneity Govern Binary Quantization of Contrastive Embeddings
Wenxuan Xiao
Binary quantization (BQ) compresses high-dimensional embeddings into one or two bits per coordinate, enabling nearest neighbor search at extreme speed. Yet a striking puzzle persis…
QuIVer: Rethinking ANN Graph Topology via Training-Free Binary Quantization
Wenxuan Xiao, Zhiyou Wang, Chengcheng Li
Approximate nearest neighbor (ANN) graph indices such as HNSW and Vamana construct their edge topology in full-precision or high-fidelity quantized metric spaces, relegating binary…