From the 1 of 17 linked papers with an AI index.
6 papers · 1 filter
Full-Stack FP4: Stable LLM Pretraining with Quantized Projections, Optimizers, and Attention
Siyu Ding, Mingchuan Ma, Jiabo Tong +3
Recent NVFP4 pretraining work has primarily optimized Transformer linear projections, leaving persistent optimizer states, optimizer computation, and low-precision attention forwar…
Trust Region On-Policy Distillation
Xingrun Xing, Haoqing Wang, Boyan Gao +2
On-Policy Distillation (OPD) is a fundamental technique for efficient post-training of large language models (LLMs), with broad applications in agent learning, multi-task enhanceme…
SpikeLLM: Scaling up Spiking Neural Network to Large Language Models via Saliency-based Spiking
Xingrun Xing, Boyan Gao, Zheng Zhang +5
Recent advancements in large language models (LLMs) with billions of parameters have improved performance in various applications, but their inference processes demand significant…
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models
Xingrun Xing, Zheng Liu, Shitao Xiao +6
Modern large language models (LLMs) driven by scaling laws, achieve intelligence emergency in large model sizes. Recently, the increasing concerns about cloud costs, latency, and p…
Enhancing Generalization via Sharpness-Aware Trajectory Matching for Dataset Condensation
Boyan Gao, Bo Zhao, Shreyank N Gowda +4
Dataset condensation aims to synthesize datasets with a few representative samples that can effectively represent the original datasets. This enables efficient training and produce…
BiPFT: Binary Pre-trained Foundation Transformer with Low-rank Estimation of Binarization Residual Polynomials
Xingrun Xing, Li Du, Xinyuan Wang +4
Pretrained foundation models offer substantial benefits for a wide range of downstream tasks, which can be one of the most potential techniques to access artificial general intelli…