2 papers
cs.CV2024
Lighthouse: A User-Friendly Library for Reproducible Video Moment Retrieval and Highlight Detection
Taichi Nishimura, Shota Nakada, Hokuto Munakata +1
We propose Lighthouse, a user-friendly library for reproducible video moment retrieval and highlight detection (MR-HD). Although researchers proposed various MR-HD approaches, the…
cs.MM2024
On the Audio Hallucinations in Large Audio-Video Language Models
Taichi Nishimura, Shota Nakada, Masayoshi Kondo
Large audio-video language models can generate descriptions for both video and audio. However, they sometimes ignore audio content, producing audio descriptions solely reliant on v…