11 papers
Medical-Checklist: Assessing the Comprehension of Medical Images by Multimodal Models
Bannapol Limanond, Masanori Suganuma, Takayuki Okatani
This paper introduces a new benchmark test, Medical-Checklist, for assessing medical multimodal models. The recent advancements in multimodal models have demonstrated significant p…
When Does High-CFG Diffusion Inversion Fail? A Controlled Study of Prompt--Latent Interactions
Yan Zeng, Yusuke Hosoya, Huyen T. T. Tran +1
Text-guided diffusion inversion is central to image editing, where an image is mapped to an initial latent and then edited by replaying the denoising process under a modified promp…
Inverting the Generation Process of Denoising Diffusion Implicit Models: Empirical Evaluation and a Novel Method
Yan Zeng, Masanori Suganuma, Takayuki Okatani
This paper studies the problem of inverting the DDIM image generation process to recover latent variables, particularly the initial noise map, from a generated image. Existing meth…
TB-Bench: Training and Testing Multi-Modal AI for Understanding Spatio-Temporal Traffic Behaviors from Dashcam Images/Videos
Korawat Charoenpitaks, Van-Quang Nguyen, Masanori Suganuma +4
The application of Multi-modal Large Language Models (MLLMs) in Autonomous Driving (AD) faces significant challenges due to their limited training on traffic-specific data and the…
RP-SLAM: Real-time Photorealistic SLAM with Efficient 3D Gaussian Splatting
Lizhi Bai, Chunqi Tian, Jun Yang +3
3D Gaussian Splatting has emerged as a promising technique for high-quality 3D rendering, leading to increasing interest in integrating 3DGS into realism SLAM systems. However, exi…
Rethinking Annotation for Object Detection: Is Annotating Small-size Instances Worth Its Cost?
Yusuke Hosoya, Masanori Suganuma, Takayuki Okatani
Detecting objects occupying only small areas in an image is difficult, even for humans. Therefore, annotating small-size object instances is hard and thus costly. This study questi…