From the 1 of 4 linked papers with an AI index.
4 papers
Towards Vision-Free CIR: Attribute-Augmented Scoring and LLM-Based Reranking for Zero-Shot Composed Image Retrieval
Ryotaro Shimada, Yu-Chieh Lin, Yuji Nozawa +3
The paper proposes a vision‑free framework for composed image retrieval that uses attribute‑augmented scoring to recover visual details and a large language model for reranking to…
On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training
Xueqing Wu, Yu-Chi Lin, Kai-Wei Chang +1
Post-training has greatly improved reasoning in frontier vision-language models, yet its gains for perception remain comparatively limited, creating a bottleneck for end-to-end vis…
CIRCLED: A Multi-turn CIR Dataset with Consistent Dialogues across Domains
Tomohisa Takeda, Yu-Chieh Lin, Yuji Nozawa +3
Existing Multi-Turn Composed Image Retrieval (MTCIR) datasets lack dialogue-historyconsistency and are restricted to the fashion domain. To address these limitations, we construct…
Prompt-Guided Attention Head Selection for Focus-Oriented Image Retrieval
Yuji Nozawa, Yu-Chieh Lin, Kazumoto Nakamura +1
The goal of this paper is to enhance pretrained Vision Transformer (ViT) models for focus-oriented image retrieval with visual prompting. In real-world image retrieval scenarios, b…