1 citations · 1 across the 3 of their papers we have counts for
4 papers
Metropolis-Hastings Captioning Game: Knowledge Fusion of Vision Language Models via Decentralized Bayesian Inference
Yuta Matsui, Ryosuke Yamaki, Ryo Ueda +2
We propose the Metropolis-Hastings Captioning Game (MHCG), a method to fuse knowledge of multiple vision-language models (VLMs) by learning from each other. Although existing metho…
Pointing out Human Answer Mistakes in a Goal-Oriented Visual Dialogue
Ryosuke Oshima, Seitaro Shinagawa, Hideki Tsunashima +2
Effective communication between humans and intelligent agents has promising applications for solving complex problems. One such approach is visual dialogue, which leverages multimo…
Community-Driven Comprehensive Scientific Paper Summarization: Insight from cvpaper.challenge
Shintaro Yamamoto, Hirokatsu Kataoka, Ryota Suzuki +2
The present paper introduces a group activity involving writing summaries of conference proceedings by volunteer participants. The rapid increase in scientific papers is a heavy bu…
Interactive Image Manipulation with Natural Language Instruction Commands
Seitaro Shinagawa, Koichiro Yoshino, Sakriani Sakti +2
We propose an interactive image-manipulation system with natural language instruction, which can generate a target image from a source image and an instruction that describes the d…