4 citations · 5 across the 5 of their papers we have counts for
7 papers
Rationale-Enhanced Decoding for Multi-modal Chain-of-Thought
Shin'ya Yamaguchi, Kosuke Nishida, Daiki Chijiwa
Large vision-language models (LVLMs) have demonstrated remarkable capabilities by integrating pre-trained vision encoders with large language models (LLMs). Similar to single-modal…
Post-pre-training for Modality Alignment in Vision-Language Foundation Models
Shin'ya Yamaguchi, Dewei Feng, Sekitoshi Kanai +2
Contrastive language image pre-training (CLIP) is an essential component of building modern vision-language foundation models. While CLIP demonstrates remarkable zero-shot performa…
Zero-shot Concept Bottleneck Models
Shin'ya Yamaguchi, Kosuke Nishida, Daiki Chijiwa +1
Concept bottleneck models (CBMs) are inherently interpretable and intervenable neural network models, which explain their final label prediction by the intermediate prediction of h…
Evaluating Time-Series Training Dataset through Lens of Spectrum in Deep State Space Models
Sekitoshi Kanai, Yasutoshi Ida, Kazuki Adachi +3
This study investigates a method to evaluate time-series datasets in terms of the performance of deep neural networks (DNNs) with state space models (deep SSMs) trained on the data…
F-Drop&Match: GANs with a Dead Zone in the High-Frequency Domain
Shin'ya Yamaguchi, Sekitoshi Kanai
Generative adversarial networks built from deep convolutional neural networks (GANs) lack the ability to exactly replicate the high-frequency components of natural images. To allev…
Constraining Logits by Bounded Function for Adversarial Robustness
Sekitoshi Kanai, Masanori Yamada, Shin'ya Yamaguchi +2
We propose a method for improving adversarial robustness by addition of a new bounded function just before softmax. Recent studies hypothesize that small logits (inputs of softmax)…