5 papers
TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework
Chao Zhang, Yuhao Wang, Derong Xu +9
Retrieval-Augmented Generation (RAG) utilizes external knowledge to augment Large Language Models' (LLMs) reliability. For flexibility, agentic RAG employs autonomous, multi-round…
From Image to Video, what do we need in multimodal LLMs?
Suyuan Huang, Haoxin Zhang, Linqing Zhong +4
Covering from Image LLMs to the more complex Video LLMs, the Multimodal Large Language Models (MLLMs) have demonstrated profound capabilities in comprehending cross-modal informati…
NoteLLM-2: Multimodal Large Representation Models for Recommendation
Chao Zhang, Haoxin Zhang, Shiwei Wu +6
Large Language Models (LLMs) have demonstrated exceptional proficiency in text understanding and embedding tasks. However, their potential in multimodal representation, particularl…
ScalingNote: Scaling up Retrievers with Large Language Models for Real-World Dense Retrieval
Suyuan Huang, Chao Zhang, Yuanyuan Wu +12
Dense retrieval in most industries employs dual-tower architectures to retrieve query-relevant documents. Due to online deployment requirements, existing real-world dense retrieval…
Benchmarking Large Language Models for Conversational Question Answering in Multi-instructional Documents
Shiwei Wu, Chen Zhang, Yan Gao +4
Instructional documents are rich sources of knowledge for completing various tasks, yet their unique challenges in conversational question answering (CQA) have not been thoroughly…