activity
20142022
most citedRegionCLIP: Region-based Language-Image Pretraining

13 citations · 18 across the 5 of their papers we have counts for

collaborators

9 papers

cs.CV202415 cited

BioDrone: A Bionic Drone-based Single Object Tracking Benchmark for Robust Vision

Xin Zhao, Shiyu Hu, Yipei Wang +8

Single object tracking (SOT) is a fundamental problem in computer vision, with a wide range of applications, including autonomous driving, augmented reality, and robot navigation.…

cs.CV2023

Eventful Transformers: Leveraging Temporal Redundancy in Vision Transformers

Matthew Dutson, Yin Li, Mohit Gupta

Vision Transformers achieve impressive accuracy across a range of visual recognition tasks. Unfortunately, their accuracy frequently comes with high computational costs. This is a…

cs.CV2023

SimHaze: game engine simulated data for real-world dehazing

Zhengyang Lou, Huan Xu, Fangzhou Mu +7

Deep models have demonstrated recent success in single-image dehazing. Most prior methods consider fully supervised training and learn from paired clean and hazy images, where a ha…

cs.CV20231 cited

Learning Procedure-aware Video Representation from Instructional Videos and Their Narrations

Yiwu Zhong, Licheng Yu, Yang Bai +3

The abundance of instructional videos and their narrations over the Internet offers an exciting avenue for understanding procedural activities. In this work, we propose to learn vi…

cs.CV20221 cited

Robust Scene Inference under Noise-Blur Dual Corruptions

Bhavya Goyal, Jean-François Lalonde, Yin Li +1

Scene inference under low-light is a challenging problem due to severe noise in the captured images. One way to reduce noise is to use longer exposure during the capture. However,…

cs.CV20211 cited

Virtuoso: Video-based Intelligence for real-time tuning on SOCs

Jayoung Lee, PengCheng Wang, Ran Xu +5

Efficient and adaptive computer vision systems have been proposed to make computer vision tasks, such as image classification and object detection, optimized for embedded or mobile…