3 papers
cs.AI2026
Traceable Cross-Source RAG for Chinese Tibetan Medicine Question Answering
Fengxian Chen, Zhilong Tao, Jiaxuan Li +2
Retrieval-augmented generation (RAG) promises grounded question answering, yet domain settings with multiple heterogeneous knowledge bases (KBs) remain challenging. In Chinese Tibe…
cs.CV2024
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects?
Jiaxuan Li, Junwen Mo, MinhDuc Vo +2
Multimodal Large Language Models (MLLMs) have made notable advances in visual understanding, yet their abilities to recognize objects modified by specific attributes remain an open…
cs.CV2023
EVCap: Retrieval-Augmented Image Captioning with External Visual-Name Memory for Open-World Comprehension
Jiaxuan Li, Duc Minh Vo, Akihiro Sugimoto +1
Large language models (LLMs)-based image captioning has the capability of describing objects not explicitly observed in training data; yet novel objects occur frequently, necessita…