3 citations · 3 across the 9 of their papers we have counts for
4 papers · 1 filter
PolyUQuest: Verifiable Structure-Aware Web RAG over Heterogeneous Graphs
Ying Liu, Yi Ye, Quanyu Feng +5
Existing retrieval-augmented generation (RAG) systems treat web pages as flat text, losing the structural and semantic signals encoded in HTML. We present PolyUQuest, a verifiable,…
From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Long-Context LLMs
Zhanchao Xu, Haoyang Li, Qingfa Xiao +4
Existing sparse attention and KV cache compression methods for long-context LLM inference typically apply fixed sparsity patterns or uniform budgets across all attention heads, ove…
When Speed meets Accuracy: an Efficient and Effective Graph Model for Temporal Link Prediction
Haoyang Li, Yuming Xu, Yiming Li +5
Temporal link prediction in dynamic graphs is a critical task with applications in diverse domains such as social networks, recommendation systems, and e-commerce platforms. While…
A Survey on Large Language Model Acceleration based on KV Cache Management
Haoyang Li, Yiming Li, Anxin Tian +7
Large Language Models (LLMs) have revolutionized a wide range of domains such as natural language processing, computer vision, and multi-modal tasks due to their ability to compreh…