4 papers
Accelerating Vision Foundation Models with Drop-in Depthwise Convolution
Carmelo Scribano, Mohammad Mahdi, Nedyalko Prisadnikov +5
Pretrained vision foundation models deliver strong performance across tasks with limited fine-tuning. However, their Vision Transformer (ViT) backbones impose high inference costs,…
Self-supervised pretraining for an iterative image size agnostic vision transformer
Nedyalko Prisadnikov, Danda Pani Paudel, Yuqian Fu +1
Vision Transformers (ViTs) dominate self-supervised learning (SSL). While they have proven highly effective for large-scale pretraining, they are computationally inefficient and sc…
Vision encoders should be image size agnostic and task driven
Nedyalko Prisadnikov, Danda Pani Paudel, Yuqian Fu +1
This position paper argues that the next generation of vision encoders should be image size agnostic and task driven. The source of our inspiration is biological. Not a structural…
A Simple and Generalist Approach for Panoptic Segmentation
Nedyalko Prisadnikov, Wouter Van Gansbeke, Danda Pani Paudel +1
Panoptic segmentation is an important computer vision task, where the current state-of-the-art solutions require specialized components to perform well. We propose a simple general…