activity
20202026
most citedMetrics and evaluations for computational and sustainable AI efficiency

1 citations · 1 across the 2 of their papers we have counts for

collaborators

7 papers

cs.LG2026

MAPLE: MoE Adaptive Plug-and-play Layer-wise Expert allocation

Lie Li, Wen Li, Junxiao Shen +1

Sparsely-activated Mixture-of-Experts (MoE) Transformers universally fix the same number of routed experts across all layers, a convention that ignores the well-documented heteroge…

cs.AR2026

SPEAR: A System for Post-Quantization Error-Adaptive Recovery Enabling Efficient Low-Bit LLM Serving

Hongyuan Liu, Yawei Li, Zhiqiang Que +3

Efficient large language model (LLM) serving is increasingly constrained by deployment cost. Quantization is a key technique for reducing serving cost, yet even state-of-the-art 4-…

cs.LG2026

FFR: Forward-Forward Learning for Regression

Xinyang Liu, Xuanyu Liang, Shiqi Ding +4

The Forward-Forward (FF) algorithm offers a computationally efficient and biologically plausible alternative to backpropagation (BP) by training neural networks through purely loca…

cs.PF20251 cited

Metrics and evaluations for computational and sustainable AI efficiency

Hongyuan Liu, Xinyang Liu, Guosheng Hu

The rapid advancement of Artificial Intelligence (AI) has created unprecedented demands for computational power, yet methods for evaluating the performance, efficiency, and environ…

cs.CL2025

Towards Extreme Pruning of LLMs with Plug-and-Play Mixed Sparsity

Chi Xu, Gefei Zhang, Yantong Zhu +4

N:M structured pruning is essential for large language models (LLMs) because it can remove less important network weights and reduce the memory and computation requirements. Existi…

cs.CL2025

Revisiting Large Language Model Pruning using Neuron Semantic Attribution

Yizhuo Ding, Xinwei Sun, Yanwei Fu +1

Model pruning technique is vital for accelerating large language models by reducing their size and computational requirements. However, the generalizability of existing pruning met…