4 citations · 7 across the 4 of their papers we have counts for
4 papers
X-DETR: A Versatile Architecture for Instance-wise Vision-Language Tasks
Zhaowei Cai, Gukyeong Kwon, Avinash Ravichandran +4
In this paper, we study the challenging instance-wise vision-language tasks, where the free-form language is required to align with the objects instead of the whole image. To addre…
MeMOT: Multi-Object Tracking with Memory
Jiarui Cai, Mingze Xu, Wei Li +4
We propose an online tracking algorithm that performs the object detection and data association under a common framework, capable of linking objects after a long time span. This is…
Task Adaptive Parameter Sharing for Multi-Task Learning
Matthew Wallingford, Hao Li, Alessandro Achille +4
Adapting pre-trained models with broad capabilities has become standard practice for learning a wide range of downstream tasks. The typical approach of fine-tuning different models…
Learning Semantic-Aware Dynamics for Video Prediction
Xinzhu Bei, Yanchao Yang, Stefano Soatto
We propose an architecture and training scheme to predict video frames by explicitly modeling dis-occlusions and capturing the evolution of semantically consistent regions in the v…