activity
19992023
most citedHigh-Performance Neural Networks for Visual Object Classification

223 citations · 837 across the 42 of their papers we have counts for

collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV20218 cited

Spatial Dependency Networks: Neural Layers for Improved Generative Image Modeling

Đorđe Miladinović, Aleksandar Stanić, Stefan Bauer +2

How to improve generative modeling by better exploiting spatial regularities and coherence in images? We introduce a novel neural network for building image generators (decoders) a…

cs.CV20203 cited

Unsupervised Object Keypoint Learning using Local Spatial Predictability

Anand Gopalakrishnan, Sjoerd van Steenkiste, Jürgen Schmidhuber

We propose PermaKey, a novel approach to representation learning based on object keypoints. It leverages the predictability of local image regions from spatial neighborhoods to ide…

cs.CV2018

Deep Watershed Detector for Music Object Recognition

Lukas Tuggener, Ismail Elezi, Jurgen Schmidhuber +1

Optical Music Recognition (OMR) is an important and challenging area within music information retrieval, the accurate detection of music symbols in digital images is a core functio…

cs.CV2018

DeepScores -- A Dataset for Segmentation, Detection and Classification of Tiny Objects

Lukas Tuggener, Ismail Elezi, Jürgen Schmidhuber +2

We present the DeepScores dataset with the goal of advancing the state-of-the-art in small objects recognition, and by placing the question of object recognition in the context of…

cs.CV2017

Improving Speaker-Independent Lipreading with Domain-Adversarial Training

Michael Wand, Juergen Schmidhuber

We present a Lipreading system, i.e. a speech recognition system using only visual features, which uses domain-adversarial training for speaker independence. Domain-adversarial tra…

cs.CV2015155 cited

Parallel Multi-Dimensional LSTM, With Application to Fast Biomedical Volumetric Image Segmentation

Marijn F. Stollenga, Wonmin Byeon, Marcus Liwicki +1

Convolutional Neural Networks (CNNs) can be shifted across 2D images or 3D videos to segment them. They have a fixed input size and typically perceive only small local contexts of…