3 papers
cs.CV2026
RegRet: Enhancing Region-Level Retrieval in Large Multimodal Models
Xun Liang, Honghui Yang, Weihang Pan +6
Region-level retrieval aims to align user-specified image regions with relevant regions or textual descriptions, playing a crucial role in realworld applications such as e-commerce…
cs.IR2023
BookGPT: A General Framework for Book Recommendation Empowered by Large Language Model
Aakas Zhiyuli, Yanfang Chen, Xuan Zhang +1
With the continuous development and change exhibited by large language model (LLM) technology, represented by generative pretrained transformers (GPTs), many classic scenarios in v…
cs.MM2023
ChinaOpen: A Dataset for Open-world Multimodal Learning
Aozhu Chen, Ziyuan Wang, Chengbo Dong +5
This paper introduces ChinaOpen, a dataset sourced from Bilibili, a popular Chinese video-sharing website, for open-world multimodal learning. While the state-of-the-art multimodal…