2 papers
cs.CL2025
Informed Routing in LLMs: Smarter Token-Level Computation for Faster Inference
Chao Han, Yijuan Liang, Zihao Xuan +3
The deployment of large language models (LLMs) in real-world applications is increasingly limited by their high inference cost. While recent advances in dynamic token-level computa…
stat.ME2025
New Bounds and Truncation Boundaries for Importance Sampling
Yijuan Liang, Guangxin Jiang, Michael C. Fu
Importance sampling (IS) is a technique that enables statistical estimation of output performance at multiple input distributions from a single nominal input distribution. IS is co…