3 papers
cs.IR2026
H3D: Benchmarking Unsupervised Text Hashing for Fine-Grained Document Deduplication
Qianren Mao, Jiaxun Lyu, Junnan Liu +4
Document hashing provides compact representations for efficient similarity search and document deduplication, but existing studies rarely compare hashing pipelines under a unified…
cs.CL2026
XRAG: eXamining the Core -- Benchmarking Foundational Components in Advanced Retrieval-Augmented Generation
Qili Zhang, Qianren Mao, Yangyifei Luo +15
Retrieval-augmented generation (RAG) synergizes the retrieval of pertinent data with the generative capabilities of Large Language Models (LLMs), ensuring that the generated output…
cs.CL2025
Privacy-Preserving Federated Embedding Learning for Localized Retrieval-Augmented Generation
Qianren Mao, Qili Zhang, Hanwen Hao +11
Retrieval-Augmented Generation (RAG) has recently emerged as a promising solution for enhancing the accuracy and credibility of Large Language Models (LLMs), particularly in Questi…