3 papers
cs.IR2024
Distillation Matters: Empowering Sequential Recommenders to Match the Performance of Large Language Model
Yu Cui, Feng Liu, Pengbo Wang +5
Owing to their powerful semantic reasoning capabilities, Large Language Models (LLMs) have been effectively utilized as recommenders, achieving impressive performance. However, the…
cs.DC2024
INSPIRIT: Optimizing Heterogeneous Task Scheduling through Adaptive Priority in Task-based Runtime Systems
Yiqing Wang, Xiaoyan Liu, Hailong Yang +5
As modern HPC computing platforms become increasingly heterogeneous, it is challenging for programmers to fully leverage the computation power of massive parallelism offered by suc…
cs.DC2024
Minions: Accelerating Large Language Model Inference with Aggregated Speculative Execution
Siqi Wang, Hailong Yang, Xuezhu Wang +10
Large language models (LLM) have recently attracted surging interest due to their outstanding capabilities across various domains. However, enabling efficient LLM inference is chal…