6 papers
SPEAR: A System for Post-Quantization Error-Adaptive Recovery Enabling Efficient Low-Bit LLM Serving
Hongyuan Liu, Yawei Li, Zhiqiang Que +3
Efficient large language model (LLM) serving is increasingly constrained by deployment cost. Quantization is a key technique for reducing serving cost, yet even state-of-the-art 4-…
FFR: Forward-Forward Learning for Regression
Xinyang Liu, Xuanyu Liang, Shiqi Ding +4
The Forward-Forward (FF) algorithm offers a computationally efficient and biologically plausible alternative to backpropagation (BP) by training neural networks through purely loca…
ColorConceptBench: A Benchmark for Probabilistic Color-Concept Understanding in Text-to-Image Models
Chenxi Ruan, Yihan Hou, Yu Xiao +2
Text-to-image (T2I) models have advanced considerably in generating high-quality images from textual descriptions. However, their ability to associate colors with concepts remains…
Metrics and evaluations for computational and sustainable AI efficiency
Hongyuan Liu, Xinyang Liu, Guosheng Hu
The rapid advancement of Artificial Intelligence (AI) has created unprecedented demands for computational power, yet methods for evaluating the performance, efficiency, and environ…
Towards Extreme Pruning of LLMs with Plug-and-Play Mixed Sparsity
Chi Xu, Gefei Zhang, Yantong Zhu +4
N:M structured pruning is essential for large language models (LLMs) because it can remove less important network weights and reduce the memory and computation requirements. Existi…
Revisiting Large Language Model Pruning using Neuron Semantic Attribution
Yizhuo Ding, Xinwei Sun, Yanwei Fu +1
Model pruning technique is vital for accelerating large language models by reducing their size and computational requirements. However, the generalizability of existing pruning met…