8 papers
QCFuse: Query-Aware Cache Fusion via Compressed View for Efficient RAG Serving
Jianxin Yan, Wangze Ni, Zhenxin Li +8
Retrieval-augmented generation (RAG) improves large language model (LLM) answer quality by grounding generation in external evidence, but processing retrieved contexts makes the pr…
Personalized w-Event Privacy for Infinite Stream Estimation
Leilei Du, Xu Zhou, Peng Cheng +4
In applications such as event monitoring, log analysis, and video querying, -event privacy protects individual data within a sliding time window while supporting accurate stream…
FGIM: a Fast Graph-based Indexes Merging Framework for Approximate Nearest Neighbor Search
Zekai Wu, Jiabao Jin, Peng Cheng +7
As the state-of-the-art methods for high-dimensional data retrieval, Approximate Nearest Neighbor Search (ANNS) approaches with graph-based indexes have attracted increasing attent…
Infinite Stream Estimation under Personalized -Event Privacy
Leilei Du, Peng Cheng, Lei Chen +3
Streaming data collection is indispensable for stream data analysis, such as event monitoring. However, publishing these data directly leads to privacy leaks. -event privacy is…
ContextCache: Context-Aware Semantic Cache for Multi-Turn Queries in Large Language Models
Jianxin Yan, Wangze Ni, Lei Chen +4
Semantic caching significantly reduces computational costs and improves efficiency by storing and reusing large language model (LLM) responses. However, existing systems rely prima…
OSS-UAgent: An Agent-based Usability Evaluation Framework for Open Source Software
Lingkai Meng, Yu Shao, Long Yuan +6
Usability evaluation is critical to the impact and adoption of open source software (OSS), yet traditional methods relying on human evaluators suffer from high costs and limited sc…