most citedRetroLLM: Empowering Large Language Models to Retrieve Fine-grained Evidence within Generation

2 citations · 3 across the 5 of their papers we have counts for

collaborators

5 papers

cs.IR2025

HyReC: Exploring Hybrid-based Retriever for Chinese

Zunran Wang, Zheng Shenpeng, Wang Shenglan +2

Hybrid-based retrieval methods, which unify dense-vector and lexicon-based retrieval, have garnered considerable attention in the industry due to performance enhancement. However,…

cs.CL2025

Data Whisperer: Efficient Data Selection for Task-Specific LLM Fine-Tuning via Few-Shot In-Context Learning

Shaobo Wang, Xiangqi Jin, Ziming Wang +8

Fine-tuning large language models (LLMs) on task-specific data is essential for their effective deployment. As dataset sizes grow, efficiently selecting optimal subsets for trainin…

cs.CL2025

Neuro-Symbolic Query Compiler

Yuyao Zhang, Zhicheng Dou, Xiaoxi Li +5

Precise recognition of search intent in Retrieval-Augmented Generation (RAG) systems remains a challenging goal, especially under resource constraints and for complex queries with…

cs.CL20251 cited

Hierarchical Document Refinement for Long-context Retrieval-augmented Generation

Jiajie Jin, Xiaoxi Li, Guanting Dong +6

Real-world RAG applications often encounter long-context input scenarios, where redundant information and noise results in higher inference costs and reduced performance. To addres…

cs.CL20242 cited

RetroLLM: Empowering Large Language Models to Retrieve Fine-grained Evidence within Generation

Xiaoxi Li, Jiajie Jin, Yujia Zhou +4

Large language models (LLMs) exhibit remarkable generative capabilities but often suffer from hallucinations. Retrieval-augmented generation (RAG) offers an effective solution by i…