9 papers
Simile Understanding in Text-to-Image Models: An Evaluation Framework
Luecheng Wang, Shintaro Ozaki, Hidetaka Kamigaito +4
Similes provide a compact and expressive way to describe visual characteristics in text prompts. Recent text-to-image models (t2i models) can produce visually compelling outputs fr…
TextTIGER: Text-based Intelligent Generation with Entity Prompt Refinement for Text-to-Image Generation
Shintaro Ozaki, Tomoyuki Jinno, Kazuki Hayashi +6
When generating images from prompts that include specific entities, the model must retain as much entity-specific knowledge as possible. However, the number of entities is almost c…
Identifying Influential N-grams in Confidence Calibration via Regression Analysis
Shintaro Ozaki, Wataru Hashimoto, Hidetaka Kamigaito +2
While large language models (LLMs) improve performance by explicit reasoning, their responses are often overconfident, even though they include linguistic expressions demonstrating…
Diagnosing Vision Language Models' Perception by Leveraging Human Methods for Color Vision Deficiencies
Kazuki Hayashi, Shintaro Ozaki, Yusuke Sakai +2
Large-scale Vision-Language Models (LVLMs) are being deployed in real-world settings that require visual inference. As capabilities improve, applications in navigation, education,…
Understanding the Impact of Confidence in Retrieval Augmented Generation: A Case Study in the Medical Domain
Shintaro Ozaki, Yuta Kato, Siyuan Feng +8
Retrieval Augmented Generation (RAG) complements the knowledge of Large Language Models (LLMs) by leveraging external information to enhance response accuracy for queries. This app…
BQA: Body Language Question Answering Dataset for Video Large Language Models
Shintaro Ozaki, Kazuki Hayashi, Miyu Oba +3
A large part of human communication relies on nonverbal cues such as facial expressions, eye contact, and body language. Unlike language or sign language, such nonverbal communicat…