3 papers
cs.IR2025
TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework
Chao Zhang, Yuhao Wang, Derong Xu +9
Retrieval-Augmented Generation (RAG) utilizes external knowledge to augment Large Language Models' (LLMs) reliability. For flexibility, agentic RAG employs autonomous, multi-round…
cs.CV2025
From Image to Video, what do we need in multimodal LLMs?
Suyuan Huang, Haoxin Zhang, Linqing Zhong +4
Covering from Image LLMs to the more complex Video LLMs, the Multimodal Large Language Models (MLLMs) have demonstrated profound capabilities in comprehending cross-modal informati…
cs.IR2025
NoteLLM-2: Multimodal Large Representation Models for Recommendation
Chao Zhang, Haoxin Zhang, Shiwei Wu +6
Large Language Models (LLMs) have demonstrated exceptional proficiency in text understanding and embedding tasks. However, their potential in multimodal representation, particularl…