24 citations · 100 across the 25 of their papers we have counts for
8 papers · 1 filter
Contextual Explainable Video Representation: Human Perception-based Understanding
Khoa Vo, Kashu Yamazaki, Phong X. Nguyen +3
Video understanding is a growing field and a subject of intense research, which includes many interesting tasks to understanding both spatial and temporal information, e.g., action…
CLIP-TSA: CLIP-Assisted Temporal Self-Attention for Weakly-Supervised Video Anomaly Detection
Hyekang Kevin Joo, Khoa Vo, Kashu Yamazaki +1
Video anomaly detection (VAD) -- commonly formulated as a multiple-instance learning problem in a weakly-supervised manner due to its labor-intensive nature -- is a challenging pro…
VLTinT: Visual-Linguistic Transformer-in-Transformer for Coherent Video Paragraph Captioning
Kashu Yamazaki, Khoa Vo, Sang Truong +2
Video paragraph captioning aims to generate a multi-sentence description of an untrimmed video with several temporal event locations in coherent storytelling. Following the human p…
AOE-Net: Entities Interactions Modeling with Adaptive Attention Mechanism for Temporal Action Proposals Generation
Khoa Vo, Sang Truong, Kashu Yamazaki +3
Temporal action proposal generation (TAPG) is a challenging task, which requires localizing action intervals in an untrimmed video. Intuitively, we as humans, perceive an action th…
AISFormer: Amodal Instance Segmentation with Transformer
Minh Tran, Khoa Vo, Kashu Yamazaki +3
Amodal Instance Segmentation (AIS) aims to segment the region of both visible and possible occluded parts of an object instance. While Mask R-CNN-based AIS approaches have shown pr…
VLCap: Vision-Language with Contrastive Learning for Coherent Video Paragraph Captioning
Kashu Yamazaki, Sang Truong, Khoa Vo +4
In this paper, we leverage the human perceiving process, that involves vision and language interaction, to generate a coherent paragraph description of untrimmed videos. We propose…