4 citations · 5 across the 6 of their papers we have counts for
6 papers
Parallel In-context Learning for Large Vision Language Models
Shin'ya Yamaguchi, Daiki Chijiwa, Tamao Sakao +1
Large vision-language models (LVLMs) employ multi-modal in-context learning (MM-ICL) to adapt to new tasks by leveraging demonstration examples. While increasing the number of demo…
MultiModal Fine-tuning with Synthetic Captions
Shohei Enomoto, Shin'ya Yamaguchi
In this paper, we address a fundamental gap between pre-training and fine-tuning of deep neural networks: while pre-training has shifted from unimodal to multimodal learning with e…
Difference Vector Equalization for Robust Fine-tuning of Vision-Language Models
Satoshi Suzuki, Shin'ya Yamaguchi, Shoichiro Takeda +7
Contrastive pre-trained vision-language models, such as CLIP, demonstrate strong generalization abilities in zero-shot classification by leveraging embeddings extracted from image…
Lossless Vocabulary Reduction for Auto-Regressive Language Models
Daiki Chijiwa, Taku Hasegawa, Kyosuke Nishida +4
Tokenization -- the process of decomposing a given text into a sequence of subwords called tokens -- is one of the key components in the development of language models. Particularl…
On the Limitation of Diffusion Models for Synthesizing Training Datasets
Shin'ya Yamaguchi, Takuma Fukuda
Synthetic samples from diffusion models are promising for leveraging in training discriminative models as replications of real training datasets. However, we found that the synthet…
Generative Semi-supervised Learning with Meta-Optimized Synthetic Samples
Shin'ya Yamaguchi
Semi-supervised learning (SSL) is a promising approach for training deep classification models using labeled and unlabeled datasets. However, existing SSL methods rely on a large u…