3 citations · 6 across the 10 of their papers we have counts for
14 papers
SIFT-VTON: Geometric Correspondence Supervision on Cross-Attention for Virtual Try-On
Kosuke Takemoto, Takafumi Koshinaka
Diffusion-based virtual try-on methods achieve photorealistic synthesis through cross-attention mechanisms that transfer garment features to target body regions. However, these app…
HYB-VITON: A Hybrid Approach to Virtual Try-On Combining Explicit and Implicit Warping
Kosuke Takemoto, Takafumi Koshinaka
Virtual try-on systems have significant potential in e-commerce, allowing customers to visualize garments on themselves. Existing image-based methods fall into two categories: thos…
Reading Is Believing: Revisiting Language Bottleneck Models for Image Classification
Honori Udo, Takafumi Koshinaka
We revisit language bottleneck models as an approach to ensuring the explainability of deep learning models for image classification. Because of inevitable information loss incurre…
Generalized domain adaptation framework for parametric back-end in speaker recognition
Qiongqiong Wang, Koji Okabe, Kong Aik Lee +1
State-of-the-art speaker recognition systems comprise a speaker embedding front-end followed by a probabilistic linear discriminant analysis (PLDA) back-end. The effectiveness of t…
Image Captioners Sometimes Tell More Than Images They See
Honori Udo, Takafumi Koshinaka
Image captioning, a.k.a. "image-to-text," which generates descriptive text from given images, has been rapidly developing throughout the era of deep learning. To what extent is the…
Task-aware Warping Factors in Mask-based Speech Enhancement
Qiongqiong Wang, Kong Aik Lee, Takafumi Koshinaka +2
This paper proposes the use of two task-aware warping factors in mask-based speech enhancement (SE). One controls the balance between speech-maintenance and noise-removal in traini…