31 citations · 39 across the 3 of their papers we have counts for
4 papers · 1 filter
TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation
Weixi Feng, Jiachen Li, Michael Saxon +3
Video generation has many unique challenges beyond those of image generation. The temporal dimension introduces extensive possible variations across frames, over which consistency…
VMFormer: End-to-End Video Matting with Transformer
Jiachen Li, Vidit Goel, Marianna Ohanyan +3
Video matting aims to predict the alpha mattes for each frame from a given input video sequence. Recent solutions to video matting have been dominated by deep convolutional neural…
Point-to-Box Network for Accurate Object Detection via Single Point Supervision
Pengfei Chen, Xuehui Yu, Xumeng Han +7
Object detection using single point supervision has received increasing attention over the years. However, the performance gap between point supervised object detection (PSOD) and…
SeMask: Semantically Masked Transformers for Semantic Segmentation
Jitesh Jain, Anukriti Singh, Nikita Orlov +4
Finetuning a pretrained backbone in the encoder part of an image transformer network has been the traditional approach for the semantic segmentation task. However, such an approach…