2 papers
cs.LG2025
SVDq: 1.25-bit and 410x Key Cache Compression for LLM Attention
Hong Yankun, Li Xing, Zhen Hui-Ling +3
For the efficient inference of Large Language Models (LLMs), the effective compression of key-value (KV) cache is essential. Three main types of KV cache compression techniques, na…
cs.LG2025
Certifying Language Model Robustness with Fuzzed Randomized Smoothing: An Efficient Defense Against Backdoor Attacks
Bowei He, Lihao Yin, Hui-Ling Zhen +4
The widespread deployment of pre-trained language models (PLMs) has exposed them to textual backdoor attacks, particularly those planted during the pre-training stage. These attack…