2 citations · 2 across the 3 of their papers we have counts for
3 papers
LARE: Latent Augmentation using Regional Embedding with Vision-Language Model
Kosuke Sakurai, Tatsuya Ishii, Ryotaro Shimizu +2
In recent years, considerable research has been conducted on vision-language models that handle both image and text data; these models are being applied to diverse downstream tasks…
Partial Visual-Semantic Embedding: Fashion Intelligence System with Sensitive Part-by-Part Learning
Ryotaro Shimizu, Takuma Nakamura, Masayuki Goto
In this study, we propose a technology called the Fashion Intelligence System based on the visual-semantic embedding (VSE) model to quantify abstract and complex expressions unique…
Fashion-Specific Attributes Interpretation via Dual Gaussian Visual-Semantic Embedding
Ryotaro Shimizu, Masanari Kimura, Masayuki Goto
Several techniques to map various types of components, such as words, attributes, and images, into the embedded space have been studied. Most of them estimate the embedded represen…