collaborators

8 papers

cs.CL2026

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction

Ziqiang Cui, Han Shi, Bowei He +8

Multi-Token Prediction (MTP) has emerged as an effective paradigm that augments a shared Large Language Model backbone with auxiliary heads, training the model to predict several f…

cs.CL2026

PSD: Pushing the Pareto Frontier of Diffusion LLMs via Parallel Speculative Decoding

Shengyin Sun, Yiming Li, Renxi Liu +7

Diffusion large language models (dLLMs) generate text by iteratively denoising masked token sequences. Although dLLMs can predict all masked positions in parallel within each step,…

cs.CL2026

Revisiting Judge Decoding from First Principles via Training-Free Distributional Divergence

Shengyin Sun, Yiming Li, Renxi Liu +5

Judge Decoding accelerates LLM inference by relaxing the strict verification of Speculative Decoding, yet it typically relies on expensive and noisy supervision. In this work, we r…

cs.LG2025

Enhanced Pre-training of Graph Neural Networks for Million-Scale Heterogeneous Graphs

Shengyin Sun, Chen Ma, Jiehao Chen

In recent years, graph neural networks (GNNs) have facilitated the development of graph data mining. However, training GNNs requires sufficient labeled task-specific data, which is…

cs.CL2025

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling

Shengyin Sun, Yiming Li, Xing Li +8

Test-time scaling has emerged as a powerful paradigm for enhancing the reasoning capabilities of large language models (LLMs) by allocating additional computational resources durin…

cs.IR2025

Hyperbolic Contrastive Learning with Model-augmentation for Knowledge-aware Recommendation

Shengyin Sun, Chen Ma

Benefiting from the effectiveness of graph neural networks (GNNs) and contrastive learning, GNN-based contrastive learning has become mainstream for knowledge-aware recommendation.…