2 papers
cs.LG2024
Towards Accurate and Efficient Sub-8-Bit Integer Training
Wenjin Guo, Donglai Liu, Weiying Xie +7
Neural network training is a memory- and compute-intensive task. Quantization, which enables low-bitwidth formats in training, can significantly mitigate the workload. To reduce qu…
cs.IR2024
Efficient and Effective Retrieval of Dense-Sparse Hybrid Vectors using Graph-based Approximate Nearest Neighbor Search
Haoyu Zhang, Jun Liu, Zhenhua Zhu +5
ANNS for embedded vector representations of texts is commonly used in information retrieval, with two important information representations being sparse and dense vectors. While it…