activity
20202026
most citedAISFormer: Amodal Instance Segmentation with Transformer

24 citations · 100 across the 25 of their papers we have counts for

collaborators
Showing 2022Show all

8 papers · 1 filter

cs.CV2022★ 1 cited

Contextual Explainable Video Representation: Human Perception-based Understanding

Khoa Vo, Kashu Yamazaki, Phong X. Nguyen +3

Video understanding is a growing field and a subject of intense research, which includes many interesting tasks to understanding both spatial and temporal information, e.g., action…

cs.CV2022★ 7 cited

CLIP-TSA: CLIP-Assisted Temporal Self-Attention for Weakly-Supervised Video Anomaly Detection

Hyekang Kevin Joo, Khoa Vo, Kashu Yamazaki +1

Video anomaly detection (VAD) -- commonly formulated as a multiple-instance learning problem in a weakly-supervised manner due to its labor-intensive nature -- is a challenging pro…

cs.CV2022★ 1 cited

VLTinT: Visual-Linguistic Transformer-in-Transformer for Coherent Video Paragraph Captioning

Kashu Yamazaki, Khoa Vo, Sang Truong +2

Video paragraph captioning aims to generate a multi-sentence description of an untrimmed video with several temporal event locations in coherent storytelling. Following the human p…

cs.CV2022★ 1 cited

AOE-Net: Entities Interactions Modeling with Adaptive Attention Mechanism for Temporal Action Proposals Generation

Khoa Vo, Sang Truong, Kashu Yamazaki +3

Temporal action proposal generation (TAPG) is a challenging task, which requires localizing action intervals in an untrimmed video. Intuitively, we as humans, perceive an action th…

cs.CV2022★ 24 cited

AISFormer: Amodal Instance Segmentation with Transformer

Minh Tran, Khoa Vo, Kashu Yamazaki +3

Amodal Instance Segmentation (AIS) aims to segment the region of both visible and possible occluded parts of an object instance. While Mask R-CNN-based AIS approaches have shown pr…

cs.CV2022★ 1 cited

VLCap: Vision-Language with Contrastive Learning for Coherent Video Paragraph Captioning

Kashu Yamazaki, Sang Truong, Khoa Vo +4

In this paper, we leverage the human perceiving process, that involves vision and language interaction, to generate a coherent paragraph description of untrimmed videos. We propose…