3 papers
cs.MM2025
Hallucination Localization in Video Captioning
Shota Nakada, Kazuhiro Saito, Yuchi Ishikawa +3
We propose a novel task, hallucination localization in video captioning, which aims to identify hallucinations in video captions at the span level (i.e. individual words or phrases…
cs.CV2025
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos
Yuchi Ishikawa, Shota Nakada, Hokuto Munakata +3
In this paper, we propose Language-Guided Contrastive Audio-Visual Masked Autoencoders (LG-CAV-MAE) to improve audio-visual representation learning. LG-CAV-MAE integrates a pretrai…
cs.MM2024
DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information
Shota Nakada, Taichi Nishimura, Hokuto Munakata +2
Current audio-visual representation learning can capture rough object categories (e.g., ``animals'' and ``instruments''), but it lacks the ability to recognize fine-grained details…