3 papers
cs.CV2026
Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA
Zhou Du, Hamid Krim, Xiao Wu +3
Recent advances in video multimodal models have significantly improved VideoQA performance. However, these systems often rely on spurious statistical correlations rather than answe…
cs.CV2025
Shot2Tactic-Caption: Multi-Scale Captioning of Badminton Videos for Tactical Understanding
Ning Ding, Keisuke Fujii, Toru Tamaki
Tactical understanding in badminton involves interpreting not only individual actions but also how tactics are dynamically executed over time. In this paper, we propose \textbf{Sho…
cs.CV2025
VIFSS: View-Invariant and Figure Skating-Specific Pose Representation Learning for Temporal Action Segmentation
Ryota Tanaka, Tomohiro Suzuki, Keisuke Fujii
Understanding human actions from videos plays a critical role across various domains, including sports analytics. In figure skating, accurately recognizing the type and timing of j…