4 papers
Zero-shot Concept Bottleneck Models
Shin'ya Yamaguchi, Kosuke Nishida, Daiki Chijiwa +1
Concept bottleneck models (CBMs) are inherently interpretable and intervenable neural network models, which explain their final label prediction by the intermediate prediction of h…
Rationale-Enhanced Decoding for Multi-modal Chain-of-Thought
Shin'ya Yamaguchi, Kosuke Nishida, Daiki Chijiwa
Large vision-language models (LVLMs) have demonstrated remarkable capabilities by integrating pre-trained vision encoders with large language models (LLMs). Similar to single-modal…
Post-pre-training for Modality Alignment in Vision-Language Foundation Models
Shin'ya Yamaguchi, Dewei Feng, Sekitoshi Kanai +2
Contrastive language image pre-training (CLIP) is an essential component of building modern vision-language foundation models. While CLIP demonstrates remarkable zero-shot performa…
Transfer Learning with Pre-trained Conditional Generative Models
Shin'ya Yamaguchi, Sekitoshi Kanai, Atsutoshi Kumagai +2
Transfer learning is crucial in training deep neural networks on new target tasks. Current transfer learning methods always assume at least one of (i) source and target task label…