4 papers
Hallucination Localization in Video Captioning
Shota Nakada, Kazuhiro Saito, Yuchi Ishikawa +3
We propose a novel task, hallucination localization in video captioning, which aims to identify hallucinations in video captions at the span level (i.e. individual words or phrases…
Listening without Looking: Modality Bias in Audio-Visual Captioning
Yuchi Ishikawa, Toranosuke Manabe, Tatsuya Komatsu +1
Audio-visual captioning aims to generate holistic scene descriptions by jointly modeling sound and vision. While recent methods have improved performance through sophisticated moda…
ProLAP: Probabilistic Language-Audio Pre-Training
Toranosuke Manabe, Yuchi Ishikawa, Hokuto Munakata +1
Language-audio joint representation learning frameworks typically depend on deterministic embeddings, assuming a one-to-one correspondence between audio and text. In real-world set…
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos
Yuchi Ishikawa, Shota Nakada, Hokuto Munakata +3
In this paper, we propose Language-Guided Contrastive Audio-Visual Masked Autoencoders (LG-CAV-MAE) to improve audio-visual representation learning. LG-CAV-MAE integrates a pretrai…