2 papers
cs.IR2026
Position-Aware Drafting for Inference Acceleration in LLM-Based Generative List-Wise Recommendation
Jiaju Chen, Chongming Gao, Chenxiao Fan +4
Large language model (LLM)-based generative list-wise recommendation has advanced rapidly, but decoding remains sequential and thus latency-prone. To accelerate inference without c…
cs.CL2026
Medical Reasoning with Large Language Models: A Survey and MR-Bench
Xiaohan Ren, Chenxiao Fan, Wenyin Ma +4
Large language models (LLMs) have achieved strong performance on medical exam-style tasks, motivating growing interest in their deployment in real-world clinical settings. However,…