collaborators

6 papers

cs.CL2026

NanoSpec: Accelerating Speculative Decoding using Minimalist In-Context Vocabularies

Zhiyang Chen, Daliang Xu, Yinyuan Zhang +3

The massive vocabulary sizes of large language models, often exceeding 100k tokens, impose a computational bottleneck on the final linear projection layer during speculative decodi…

cs.AI2026

Teaching AI Through Benchmark Construction: QuestBench as a Course-Based Practice for Accountable Knowledge Work

Haiyang Shen, Jiuzheng Wang, Taian Guo +9

As AI becomes part of everyday learning, many courses teach students to use it mainly as a productivity tool: how to prompt, search, summarize, write, code, and use tools more effi…

cs.LG2026

The Last Human-Written Paper: Agent-Native Research Artifacts

Jiachen Liu, Jiaxin Pei, Jintao Huang +34

Scientific publication compresses a branching, iterative research process into a linear narrative, discarding the majority of what was discovered along the way. This compilation im…

cs.AI2026

DRAGON: Domain-specific Robust Automatic Data Generation for RAG Optimization

Haiyang Shen, Hang Yan, Zhongshi Xing +6

Retrieval-augmented generation (RAG) can substantially enhance the performance of LLMs on knowledge-intensive tasks. Various RAG paradigms - including vanilla, planning-based, and…

cs.SE2026

A First Look at Bugs in LLM Inference Engines

Mugeng Liu, Siqi Zhong, Weichen Bi +5

Large language model-specific inference engines (in short as \emph{LLM inference engines}) have become a fundamental component of modern AI infrastructure, enabling the deployment…

cs.CL2025

Accelerating Mobile Language Model via Speculative Decoding and NPU-Coordinated Execution

Zhiyang Chen, Daliang Xu, Haiyang Shen +5

Performing Retrieval-Augmented Generation (RAG) directly on mobile devices is promising for data privacy and responsiveness but is hindered by the architectural constraints of mobi…