24 citations · 32 across the 3 of their papers we have counts for
3 papers
cs.CV2022★ 7 cited
CLIP-TSA: CLIP-Assisted Temporal Self-Attention for Weakly-Supervised Video Anomaly Detection
Hyekang Kevin Joo, Khoa Vo, Kashu Yamazaki +1
Video anomaly detection (VAD) -- commonly formulated as a multiple-instance learning problem in a weakly-supervised manner due to its labor-intensive nature -- is a challenging pro…
cs.CV2022★ 1 cited
VLTinT: Visual-Linguistic Transformer-in-Transformer for Coherent Video Paragraph Captioning
Kashu Yamazaki, Khoa Vo, Sang Truong +2
Video paragraph captioning aims to generate a multi-sentence description of an untrimmed video with several temporal event locations in coherent storytelling. Following the human p…
cs.CV2022★ 24 cited
AISFormer: Amodal Instance Segmentation with Transformer
Minh Tran, Khoa Vo, Kashu Yamazaki +3
Amodal Instance Segmentation (AIS) aims to segment the region of both visible and possible occluded parts of an object instance. While Mask R-CNN-based AIS approaches have shown pr…