9 papers
ProServe: Unified Multi-Priority Request Scheduling for LLM Serving
Weizhe Huang, Tao Peng, Tongxuan Liu +4
The widespread deployment of large language models (LLMs) for interactive applications necessitates serving systems that can handle thousands of concurrent requests with diverse Se…
RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference
Ben Wan, Yan Feng, Zihan Tang +4
DeepSeek-OCR leverages visual-text compression to reduce long-text processing costs and accelerate inference, yet visual tokens remain prone to redundant textual and structural inf…
TARAC: Mitigating Hallucination in LVLMs via Temporal Attention Real-time Accumulative Connection
Lei Jiang, Chunzhao Xie, Tongxuan Liu +6
Large Vision-Language Models have demonstrated remarkable capabilities, yet they suffer from hallucinations that limit practical deployment. While various mitigation strategies exi…
Survey of Computerized Adaptive Testing: A Machine Learning Perspective
Yan Zhuang, Qi Liu, Haoyang Bi +12
Computerized Adaptive Testing (CAT) offers an efficient and personalized method for assessing examinee proficiency by dynamically adjusting test questions based on individual perfo…
xLLM Technical Report
Tongxuan Liu, Tao Peng, Peijun Yang +50
We introduce xLLM, an intelligent and efficient Large Language Model (LLM) inference framework designed for high-performance, large-scale enterprise-grade serving, with deep optimi…
IFDNS: An Iterative Feedback-Driven Neuro-Symbolic Method for Faithful Logical Reasoning
Xiaoheng Wang, Tongxuan Liu, Zi Gong +5
Large language models (LLMs) have demonstrated impressive capabilities across a wide range of reasoning tasks, including logical and mathematical problem-solving. While prompt-base…