3 papers
cs.CV2025
Object-Centric Framework for Video Moment Retrieval
Zongyao Li, Yongkang Wong, Satoshi Yamazaki +2
Most existing video moment retrieval methods rely on temporal sequences of frame- or clip-level features that primarily encode global visual and semantic information. However, such…
cs.CV2025
KFS-Bench: Comprehensive Evaluation of Key Frame Sampling in Long Video Understanding
Zongyao Li, Kengo Ishida, Satoshi Yamazaki +2
We propose KFS-Bench, the first benchmark for key frame sampling in long video question answering (QA), featuring multi-scene annotations to enable direct and robust evaluation of…
cs.CV2024
Enhanced Data Transfer Cooperating with Artificial Triplets for Scene Graph Generation
KuanChao Chu, Satoshi Yamazaki, Hideki Nakayama
This work focuses on training dataset enhancement of informative relational triplets for Scene Graph Generation (SGG). Due to the lack of effective supervision, the current SGG mod…