From the 1 of 10 linked papers with an AI index.
10 papers
Medical-Checklist: Assessing the Comprehension of Medical Images by Multimodal Models
Bannapol Limanond, Masanori Suganuma, Takayuki Okatani
This paper introduces a new benchmark test, Medical-Checklist, for assessing medical multimodal models. The recent advancements in multimodal models have demonstrated significant p…
Cascaded Multi-Scale Attention for Enhanced Multi-Scale Feature Extraction and Interaction with Low-Resolution Images
Xiangyong Lu, Masanori Suganuma, Takayuki Okatani
The paper introduces Cascaded Multi-Scale Attention (CMSA), an attention module for CNN‑ViT hybrid networks that extracts and fuses multi‑scale features without downsampling, impro…
When Does High-CFG Diffusion Inversion Fail? A Controlled Study of Prompt--Latent Interactions
Yan Zeng, Yusuke Hosoya, Huyen T. T. Tran +1
Text-guided diffusion inversion is central to image editing, where an image is mapped to an initial latent and then edited by replaying the denoising process under a modified promp…
Inverting the Generation Process of Denoising Diffusion Implicit Models: Empirical Evaluation and a Novel Method
Yan Zeng, Masanori Suganuma, Takayuki Okatani
This paper studies the problem of inverting the DDIM image generation process to recover latent variables, particularly the initial noise map, from a generated image. Existing meth…
An Improved Method for Personalizing Diffusion Models
Yan Zeng, Masanori Suganuma, Takayuki Okatani
Diffusion models have demonstrated impressive image generation capabilities. Personalized approaches, such as textual inversion and Dreambooth, enhance model individualization usin…
TB-Bench: Training and Testing Multi-Modal AI for Understanding Spatio-Temporal Traffic Behaviors from Dashcam Images/Videos
Korawat Charoenpitaks, Van-Quang Nguyen, Masanori Suganuma +4
The application of Multi-modal Large Language Models (MLLMs) in Autonomous Driving (AD) faces significant challenges due to their limited training on traffic-specific data and the…