2 papers
cs.MM2024
On the Audio Hallucinations in Large Audio-Video Language Models
Taichi Nishimura, Shota Nakada, Masayoshi Kondo
Large audio-video language models can generate descriptions for both video and audio. However, they sometimes ignore audio content, producing audio descriptions solely reliant on v…
cs.CV2023
Leveraging Image-Text Similarity and Caption Modification for the DataComp Challenge: Filtering Track and BYOD Track
Shuhei Yokoo, Peifei Zhu, Yuchi Ishikawa +3
Large web crawl datasets have already played an important role in learning multimodal features with high generalization capabilities. However, there are still very limited studies…