2 papers
cs.LG2026
OffQ: Taming Structured Outliers in LLM Quantization by Offsetting
Haoqi Wang, Lorenz K. Mueller, Jiawei Zhuang +2
Low-bit quantization has been widely adopted to accelerate the inference of large language models (LLMs) by significantly reducing computational cost and memory usage. However, act…
cs.LG2024
OpenOOD v1.5: Enhanced Benchmark for Out-of-Distribution Detection
Jingyang Zhang, Jingkang Yang, Pengyun Wang +9
Out-of-Distribution (OOD) detection is critical for the reliable operation of open-world intelligent systems. Despite the emergence of an increasing number of OOD detection methods…