2 papers
cs.CV2025
Valley: Video Assistant with Large Language model Enhanced abilitY
Ruipu Luo, Ziwang Zhao, Min Yang +6
Large Language Models (LLMs), with remarkable conversational capability, have emerged as AI assistants that can handle both visual and textual modalities. However, their effectiven…
cs.CV2024
Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering
Junxiao Xue, Quan Deng, Fei Yu +3
Multimodal large language models (MLLMs), such as GPT-4o, Gemini, LLaVA, and Flamingo, have made significant progress in integrating visual and textual modalities, excelling in tas…