collaborators

5 papers

eess.AS2025

CASTELLA: Long Audio Dataset with Captions and Temporal Boundaries

Hokuto Munakata, Takehiro Imamura, Taichi Nishimura +1

We introduce CASTELLA, a human-annotated audio benchmark for the task of audio moment retrieval (AMR). Although AMR has various useful potential applications, there is still no est…

cs.MM2025

Hallucination Localization in Video Captioning

Shota Nakada, Kazuhiro Saito, Yuchi Ishikawa +3

We propose a novel task, hallucination localization in video captioning, which aims to identify hallucinations in video captions at the span level (i.e. individual words or phrases…

eess.AS2025

ProLAP: Probabilistic Language-Audio Pre-Training

Toranosuke Manabe, Yuchi Ishikawa, Hokuto Munakata +1

Language-audio joint representation learning frameworks typically depend on deterministic embeddings, assuming a one-to-one correspondence between audio and text. In real-world set…

cs.CV2025

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos

Yuchi Ishikawa, Shota Nakada, Hokuto Munakata +3

In this paper, we propose Language-Guided Contrastive Audio-Visual Masked Autoencoders (LG-CAV-MAE) to improve audio-visual representation learning. LG-CAV-MAE integrates a pretrai…

cs.MM2024

DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information

Shota Nakada, Taichi Nishimura, Hokuto Munakata +2

Current audio-visual representation learning can capture rough object categories (e.g., ``animals'' and ``instruments''), but it lacks the ability to recognize fine-grained details…