Showing cs.CVShow all
3 papers · 1 filter
cs.CV2024
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy
Te Yang, Jian Jia, Xiangyu Zhu +9
Large Language Models (LLMs) have strong instruction-following capability to interpret and execute tasks as directed by human commands. Multimodal Large Language Models (MLLMs) hav…
cs.CV2024
Spatiotemporal Graph Guided Multi-modal Network for Livestreaming Product Retrieval
Xiaowan Hu, Yiyi Chen, Yan Li +5
With the rapid expansion of e-commerce, more consumers have become accustomed to making purchases via livestreaming. Accurately identifying the products being sold by salespeople,…
cs.CV2024
Knowledge Condensation and Reasoning for Knowledge-based VQA
Dongze Hao, Jian Jia, Longteng Guo +8
Knowledge-based visual question answering (KB-VQA) is a challenging task, which requires the model to leverage external knowledge for comprehending and answering questions grounded…