76 citations · 98 across the 2 of their papers we have counts for
2 papers
cs.CV2022★ 76 cited
SegViT: Semantic Segmentation with Plain Vision Transformers
Bowen Zhang, Zhi Tian, Quan Tang +4
We explore the capability of plain Vision Transformers (ViTs) for semantic segmentation and propose the SegVit. Previous ViT-based segmentation networks usually learn a pixel-level…
cs.CV2022★ 22 cited
Context-Driven Detection of Invertebrate Species in Deep-Sea Video
R. Austin McEver, Bowen Zhang, Connor Levenson +2
Each year, underwater remotely operated vehicles (ROVs) collect thousands of hours of video of unexplored ocean habitats revealing a plethora of information regarding biodiversity…