works on

From the 1 of 17 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2026

Full-Stack FP4: Stable LLM Pretraining with Quantized Projections, Optimizers, and Attention

Siyu Ding, Mingchuan Ma, Jiabo Tong +3

Recent NVFP4 pretraining work has primarily optimized Transformer linear projections, leaving persistent optimizer states, optimizer computation, and low-precision attention forwar…

cs.LG2026

Trust Region On-Policy Distillation

Xingrun Xing, Haoqing Wang, Boyan Gao +2

On-Policy Distillation (OPD) is a fundamental technique for efficient post-training of large language models (LLMs), with broad applications in agent learning, multi-task enhanceme…

cs.LG2025

SpikeLLM: Scaling up Spiking Neural Network to Large Language Models via Saliency-based Spiking

Xingrun Xing, Boyan Gao, Zheng Zhang +5

Recent advancements in large language models (LLMs) with billions of parameters have improved performance in various applications, but their inference processes demand significant…

cs.LG2025

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models

Xingrun Xing, Zheng Liu, Shitao Xiao +6

Modern large language models (LLMs) driven by scaling laws, achieve intelligence emergency in large model sizes. Recently, the increasing concerns about cloud costs, latency, and p…

cs.LG2025

Enhancing Generalization via Sharpness-Aware Trajectory Matching for Dataset Condensation

Boyan Gao, Bo Zhao, Shreyank N Gowda +4

Dataset condensation aims to synthesize datasets with a few representative samples that can effectively represent the original datasets. This enables efficient training and produce…

cs.LG2024

BiPFT: Binary Pre-trained Foundation Transformer with Low-rank Estimation of Binarization Residual Polynomials

Xingrun Xing, Li Du, Xinyuan Wang +4

Pretrained foundation models offer substantial benefits for a wide range of downstream tasks, which can be one of the most potential techniques to access artificial general intelli…