3 citations · 3 across the 8 of their papers we have counts for
5 papers · 1 filter
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos
Yuchi Ishikawa, Shota Nakada, Hokuto Munakata +3
In this paper, we propose Language-Guided Contrastive Audio-Visual Masked Autoencoders (LG-CAV-MAE) to improve audio-visual representation learning. LG-CAV-MAE integrates a pretrai…
Data Collection-free Masked Video Modeling
Yuchi Ishikawa, Masayoshi Kondo, Yoshimitsu Aoki
Pre-training video transformers generally requires a large amount of data, presenting significant challenges in terms of data collection costs and concerns related to privacy, lice…
Leveraging Image-Text Similarity and Caption Modification for the DataComp Challenge: Filtering Track and BYOD Track
Shuhei Yokoo, Peifei Zhu, Yuchi Ishikawa +3
Large web crawl datasets have already played an important role in learning multimodal features with high generalization capabilities. However, there are still very limited studies…
Alleviating Over-segmentation Errors by Detecting Action Boundaries
Yuchi Ishikawa, Seito Kasai, Yoshimitsu Aoki +1
We propose an effective framework for the temporal action segmentation task, namely an Action Segment Refinement Framework (ASRF). Our model architecture consists of a long-term fe…
Retrieving and Highlighting Action with Spatiotemporal Reference
Seito Kasai, Yuchi Ishikawa, Masaki Hayashi +3
In this paper, we present a framework that jointly retrieves and spatiotemporally highlights actions in videos by enhancing current deep cross-modal retrieval methods. Our work tak…