4 papers
PANICL: Mitigating Over-Reliance on Single Prompt in Visual In-Context Learning
Jiahao Zhang, Bowen Wang, Hong Liu +2
Visual In-Context Learning (VICL) uses input-output image pairs, referred to as in-context pairs (or examples), as prompts alongside query images to guide models in performing dive…
Taming the Untamed: Graph-Based Knowledge Retrieval and Reasoning for MLLMs to Conquer the Unknown
Bowen Wang, Zhouqiang Jiang, Yasuaki Susumu +3
The real value of knowledge lies not just in its accumulation, but in its potential to be harnessed effectively to conquer the unknown. Although recent multimodal large language mo…
E-InMeMo: Enhanced Prompting for Visual In-Context Learning
Jiahao Zhang, Bowen Wang, Hong Liu +3
Large-scale models trained on extensive datasets have become the standard due to their strong generalizability across diverse tasks. In-context learning (ICL), widely used in natur…
VASCAR: Content-Aware Layout Generation via Visual-Aware Self-Correction
Jiahao Zhang, Ryota Yoshihashi, Shunsuke Kitada +2
Large language models (LLMs) have proven effective for layout generation due to their ability to produce structure-description languages, such as HTML or JSON. In this paper, we ar…