activity
20232026
most citedFlattenQuant: Breaking Through the Inference Compute-bound for Large Language Models with Per-tensor Quantization

1 citations · 1 across the 4 of their papers we have counts for

collaborators

5 papers

cs.SE2026

Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research

Lezhi Yu, Xiaogang Xu, Yuhua Zhou +2

LLM agents used for scientific experimentation must do more than generate executable code: they must implement the reference method faithfully, design experiments that test the pap…

cs.LG2026

BUDDY: BUdget-Driven DYnamic Depth Routing for Adaptive Large Language Model Inference

Yuhua Zhou, Shaoqi Yu, Shichao Weng +4

Large language models (LLMs) incur high inference cost due to their depth and parameter scale. Depth pruning can reduce latency by skipping redundant Transformer blocks, but existi…

cs.LG20241 cited

FlattenQuant: Breaking Through the Inference Compute-bound for Large Language Models with Per-tensor Quantization

Yi Zhang, Fei Yang, Shuang Peng +2

Large language models (LLMs) have demonstrated state-of-the-art performance across various tasks. However, the latency of inference and the large GPU memory consumption of LLMs res…

cs.CL2023

Holmes: Towards Distributed Training Across Clusters with Heterogeneous NIC Environment

Fei Yang, Shuang Peng, Ning Sun +5

Large language models (LLMs) such as GPT-3, OPT, and LLaMA have demonstrated remarkable accuracy in a wide range of tasks. However, training these models can incur significant expe…

cs.LG2023

Exploring Post-Training Quantization of Protein Language Models

Shuang Peng, Fei Yang, Ning Sun +3

Recent advancements in unsupervised protein language models (ProteinLMs), like ESM-1b and ESM-2, have shown promise in different protein prediction tasks. However, these models fac…