2 papers
cs.CV2026
VVitCutLER: Towards Unsupervised Object Detection and Segmentation in Videos
Zhijing Lu, Khurram Azeem Hashmi, Didier Stricker +1
Unsupervised pixel-level video understanding remains challenging in real-world scenarios, where motion blur, occlusion, and fast object dynamics often cause temporal drift and flic…
cs.CV2024
GenFormer -- Generated Images are All You Need to Improve Robustness of Transformers on Small Datasets
Sven Oehri, Nikolas Ebert, Ahmed Abdullah +2
Recent studies showcase the competitive accuracy of Vision Transformers (ViTs) in relation to Convolutional Neural Networks (CNNs), along with their remarkable robustness. However,…