papers

Publications (30)

cs.CV2018

Multi-label Object Attribute Classification using a Convolutional Neural Network

Soubarna Banik, Mikko Lauri, Simone Frintrop

Objects of different classes can be described using a limited number of attributes such as color, shape, pattern, and texture. Learning to detect object attributes instead of only…

cs.CV2018

AttentionMask: Attentive, Efficient Object Proposal Generation Focusing on Small Objects

Christian Wilms, Simone Frintrop

We propose a novel approach for class-agnostic object proposal generation, which is efficient and especially well-suited to detect small objects. Efficiency is achieved by scale-sp…

cs.CV2024

High-Level Parallelism and Nested Features for Dynamic Inference Cost and Top-Down Attention

André Peter Kelm, Niels Hannemann, Bruno Heberle +5

This paper introduces a novel network topology that seamlessly integrates dynamic inference cost with a top-down attention mechanism, addressing two significant gaps in traditional…

cs.CV2021

CloudAAE: Learning 6D Object Pose Regression with On-line Data Synthesis on Point Clouds

Ge Gao, Mikko Lauri, Xiaolin Hu +2

It is often desired to train 6D pose estimation systems on synthetic data because manual annotation is expensive. However, due to the large domain gap between the synthetic and rea…

cs.CV2024

AnomalousPatchCore: Exploring the Use of Anomalous Samples in Industrial Anomaly Detection

Mykhailo Koshil, Tilman Wegener, Detlef Mentrup +2

Visual inspection, or industrial anomaly detection, is one of the most common quality control types in manufacturing. The task is to identify the presence of an anomaly given an im…

cs.RO2020

Multi-Sensor Next-Best-View Planning as Matroid-Constrained Submodular Maximization

Mikko Lauri, Joni Pajarinen, Jan Peters +1

3D scene models are useful in robotics for tasks such as path planning, object manipulation, and structural inspection. We consider the problem of creating a 3D model using depth i…

cs.CV2024

SOS: Segment Object System for Open-World Instance Segmentation With Object Priors

Christian Wilms, Tim Rolff, Maris Hillemann +2

We propose an approach for Open-World Instance Segmentation (OWIS), a task that aims to segment arbitrary unknown objects in images by generalizing from a limited set of annotated…

cs.CV2021

Superpixel-based Refinement for Object Proposal Generation

Christian Wilms, Simone Frintrop

Precise segmentation of objects is an important problem in tasks like class-agnostic object proposal generation or instance segmentation. Deep learning-based systems usually genera…

cs.CV2017

Saliency-guided Adaptive Seeding for Supervoxel Segmentation

Ge Gao, Mikko Lauri, Jianwei Zhang +1

We propose a new saliency-guided method for generating supervoxels in 3D space. Rather than using an evenly distributed spatial seeding procedure, our method uses visual saliency t…

cs.CV2022

Immersive Neural Graphics Primitives

Ke Li, Tim Rolff, Susanne Schmidt +4

Neural radiance field (NeRF), in particular its extension by instant neural graphics primitives, is a novel rendering method for view synthesis that uses real-world images to build…

cs.RO2018

Attention based visual analysis for fast grasp planning with multi-fingered robotic hand

Zhen Deng, Ge Gao, Simone Frintrop +1

We present an attention based visual analysis framework to compute grasp-relevant information in order to guide grasp planning using a multi-fingered robotic hand. Our approach use…

cs.CV2023

Small, but important: Traffic light proposals for detecting small traffic lights and beyond

Tom Sanitz, Christian Wilms, Simone Frintrop

Traffic light detection is a challenging problem in the context of self-driving cars and driver assistance systems. While most existing systems produce good results on large traffi…

cs.CV2022

Localizing Small Apples in Complex Apple Orchard Environments

Christian Wilms, Robert Johanson, Simone Frintrop

The localization of fruits is an essential first step in automated agricultural pipelines for yield estimation or fruit picking. One example of this is the localization of apples i…

cs.CV2025

Automatic Intermodal Loading Unit Identification using Computer Vision: A Scoping Review

Emre Gülsoylu, Alhassan Abdelhalim, Derya Kara Boztas +4

Background: The standardisation of Intermodal Loading Units (ILUs), including containers, semi-trailers, and swap bodies, has transformed global trade, yet efficient and robust ide…

cs.CV2025

Walk the Lines 2: Contour Tracking for Detailed Segmentation

André Peter Kelm, Max Braeschke, Emre Gülsoylu +1

This paper presents Walk the Lines 2 (WtL2), a unique contour tracking algorithm specifically adapted for detailed segmentation of infrared (IR) ships and various objects in RGB.1…

cs.CV2020

6D Object Pose Regression via Supervised Learning on Point Clouds

Ge Gao, Mikko Lauri, Yulong Wang +3

This paper addresses the task of estimating the 6 degrees of freedom pose of a known 3D object from depth information represented by a point cloud. Deep features learned by convolu…

cs.CV2025

Fusing Monocular RGB Images with AIS Data to Create a 6D Pose Estimation Dataset for Marine Vessels

Fabian Holst, Emre Gülsoylu, Simone Frintrop

The paper presents a novel technique for creating a 6D pose estimation dataset for marine vessels by fusing monocular RGB images with Automatic Identification System (AIS) data. Th…

cs.CV2022

Segmenting Medical Instruments in Minimally Invasive Surgeries using AttentionMask

Christian Wilms, Alexander Michael Gerlach, Rüdiger Schmitz +1

Precisely locating and segmenting medical instruments in images of minimally invasive surgeries, medical instrument segmentation, is an essential first step for several tasks in me…

cs.CV2018

Occlusion Resistant Object Rotation Regression from Point Cloud Segments

Ge Gao, Mikko Lauri, Jianwei Zhang +1

Rotation estimation of known rigid objects is important for robotic applications such as dexterous manipulation. Most existing methods for rotation estimation use intermediate repr…

eess.AS2023

Audio-Visual Speech Enhancement with Score-Based Generative Models

Julius Richter, Simone Frintrop, Timo Gerkmann

This paper introduces an audio-visual speech enhancement system that leverages score-based generative models, also known as diffusion models, conditioned on visual information. In…

cs.CV2023

Teacher Network Calibration Improves Cross-Quality Knowledge Distillation

Pia Čuk, Robin Senge, Mikko Lauri +1

We investigate cross-quality knowledge distillation (CQKD), a knowledge distillation method where knowledge from a teacher network trained with full-resolution images is transferre…

cs.CV2025

TRUDI and TITUS: A Multi-Perspective Dataset and A Three-Stage Recognition System for Transportation Unit Identification

Emre Gülsoylu, André Kelm, Lennart Bengtson +5

Identifying transportation units (TUs) is essential for improving the efficiency of port logistics. However, progress in this field has been hindered by the lack of publicly availa…

cs.CV2021

DeepFH Segmentations for Superpixel-based Object Proposal Refinement

Christian Wilms, Simone Frintrop

Class-agnostic object proposal generation is an important first step in many object detection pipelines. However, object proposals of modern systems are rather inaccurate in terms…

cs.CV2025

SegSLR: Promptable Video Segmentation for Isolated Sign Language Recognition

Sven Schreiber, Noha Sarhan, Simone Frintrop +1

Isolated Sign Language Recognition (ISLR) approaches primarily rely on RGB data or signer pose information. However, combining these modalities often results in the loss of crucial…

cs.CV2025

SAD: Semi-supervised Small Apple Detection in Orchard Environments

Robert Johanson, Christian Wilms, Ole Johannsen +1

Crop detection is integral for precision agriculture applications such as automated yield estimation or fruit picking. However, crop detection, e.g., apple detection in orchard env…

cs.LG2024

Select High-Level Features: Efficient Experts from a Hierarchical Classification Network

André Kelm, Niels Hannemann, Bruno Heberle +5

This study introduces a novel expert generation method that dynamically reduces task and computational complexity without compromising predictive performance. It is based on a new…

cs.CV2024

The MSR-Video to Text Dataset with Clean Annotations

Haoran Chen, Jianmin Li, Simone Frintrop +1

Video captioning automatically generates short descriptions of the video content, usually in form of a single sentence. Many methods have been proposed for solving this task. A lar…

cs.CV2025

ContextLoss: Context Information for Topology-Preserving Segmentation

Benedict Schacht, Imke Greving, Simone Frintrop +2

In image segmentation, preserving the topology of segmented structures like vessels, membranes, or roads is crucial. For instance, topological errors on road networks can significa…

cs.CV2017

Object proposal generation applying the distance dependent Chinese restaurant process

Mikko Lauri, Simone Frintrop

In application domains such as robotics, it is useful to represent the uncertainty related to the robot's belief about the state of its environment. Algorithms that only yield a si…

cs.RO2017

Multi-Robot Active Information Gathering with Periodic Communication

Mikko Lauri, Eero Heinänen, Simone Frintrop

A team of robots sharing a common goal can benefit from coordination of the activities of team members, helping the team to reach the goal more reliably or quickly. We address the…