4 papers
Towards Reliable and Holistic Visual In-Context Learning Prompt Selection
Wenxiao Wu, Jing-Hao Xue, Chengming Xu +5
Visual In-Context Learning (VICL) has emerged as a prominent approach for adapting visual foundation models to novel tasks, by effectively exploiting contextual information embedde…
VideoLucy: Deep Memory Backtracking for Long Video Understanding
Jialong Zuo, Yongtai Deng, Lingdong Kong +7
Recent studies have shown that agent-based systems leveraging large language models (LLMs) for key information retrieval and integration have emerged as a promising approach for lo…
CTR-Driven Advertising Image Generation with Multimodal Large Language Models
Xingye Chen, Wei Feng, Zhenbang Du +16
In web data, advertising images are crucial for capturing user attention and improving advertising effectiveness. Most existing methods generate background for products primarily f…
Cross-video Identity Correlating for Person Re-identification Pre-training
Jialong Zuo, Ying Nie, Hanyu Zhou +5
Recent researches have proven that pre-training on large-scale person images extracted from internet videos is an effective way in learning better representations for person re-ide…