Publications (55)
Color Constancy by GANs: An Experimental Survey
Partha Das, Anil S. Baslamisli, Yang Liu +2
In this paper, we formulate the color constancy task as an image-to-image translation problem using GANs. By conducting a large set of experiments on different datasets, an experim…
Gaussian Mapping for Evolving Scenes
Vladimir Yugay, Thies Kersten, Luca Carlone +3
Mapping systems with novel view synthesis (NVS) capabilities, most notably 3D Gaussian Splatting (3DGS), are widely used in computer vision, as well as in various applications, inc…
MAGiC-SLAM: Multi-Agent Gaussian Globally Consistent SLAM
Vladimir Yugay, Theo Gevers, Martin R. Oswald
Simultaneous localization and mapping (SLAM) systems with novel view synthesis capabilities are widely used in computer vision, with applications in augmented reality, robotics, an…
ShapeSplat: A Large-scale Dataset of Gaussian Splats and Their Self-Supervised Pretraining
Qi Ma, Yue Li, Bin Ren +5
3D Gaussian Splatting (3DGS) has become the de facto method of 3D representation in many vision tasks. This calls for the 3D understanding directly in this representation space. To…
Automatic Generation of Dense Non-rigid Optical Flow
Hoà ng-Ãn Lê, Tushar Nimbhorkar, Thomas Mensink +3
There hardly exists any large-scale datasets with dense optical flow of non-rigid motion from real-world imagery as of today. The reason lies mainly in the required setup to derive…
Kinship Identification through Joint Learning Using Kinship Verification Ensembles
Wei Wang, Shaodi You, Sezer Karaoglu +1
Kinship verification is a well-explored task: identifying whether or not two persons are kin. In contrast, kinship identification has been largely ignored so far. Kinship identific…
Joint Learning of Intrinsic Images and Semantic Segmentation
Anil S. Baslamisli, Thomas T. Groenestege, Partha Das +3
Semantic segmentation of outdoor scenes is problematic when there are variations in imaging conditions. It is known that albedo (reflectance) is invariant to all kinds of illuminat…
Deception Detection by 2D-to-3D Face Reconstruction from Videos
Minh Ngô, Burak Mandira, Selim Fırat Yılmaz +5
Lies and deception are common phenomena in society, both in our private and professional lives. However, humans are notoriously bad at accurate deception detection. Based on the li…
Gaussian-SLAM: Photo-realistic Dense SLAM with Gaussian Splatting
Vladimir Yugay, Yue Li, Theo Gevers +1
We present a dense simultaneous localization and mapping (SLAM) method that uses 3D Gaussians as a scene representation. Our approach enables interactive-time reconstruction and ph…
Road Detection by One-Class Color Classification: Dataset and Experiments
Jose M. Alvarez, Theo Gevers, Antonio M. Lopez
Detecting traversable road areas ahead a moving vehicle is a key process for modern autonomous driving systems. A common approach to road detection consists of exploiting color fea…
Grab-3D: Detecting AI-Generated Videos from 3D Geometric Temporal Consistency
Wenhan Chen, Sezer Karaoglu, Theo Gevers
Recent advances in diffusion-based generation techniques enable AI models to produce highly realistic videos, heightening the need for reliable detection mechanisms. However, exist…
SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language Pretraining
Yue Li, Qi Ma, Runyi Yang +10
Recognizing arbitrary or previously unseen categories is essential for comprehensive real-world 3D scene understanding. Currently, all existing methods rely on 2D or textual modali…
Physics-based Shading Reconstruction for Intrinsic Image Decomposition
Anil S. Baslamisli, Yang Liu, Sezer Karaoglu +1
We investigate the use of photometric invariance and deep learning to compute intrinsic images (albedo and shading). We propose albedo and shading gradient descriptors which are de…
MorphPool: Efficient Non-linear Pooling & Unpooling in CNNs
Rick Groenendijk, Leo Dorst, Theo Gevers
Pooling is essentially an operation from the field of Mathematical Morphology, with max pooling as a limited special case. The more general setting of MorphPooling greatly extends…
Spatio-temporal Features for Generalized Detection of Deepfake Videos
Ipek Ganiyusufoglu, L. Minh Ngô, Nedko Savov +2
For deepfake detection, video-level detectors have not been explored as extensively as image-level detectors, which do not exploit temporal data. In this paper, we empirically show…
T-MAE: Temporal Masked Autoencoders for Point Cloud Representation Learning
Weijie Wei, Fatemeh Karimi Nejadasl, Theo Gevers +1
The scarcity of annotated data in LiDAR point cloud understanding hinders effective representation learning. Consequently, scholars have been actively investigating efficacious sel…
TrimBot2020: an outdoor robot for automatic gardening
Nicola Strisciuglio, Radim Tylecek, Michael Blaich +9
Robots are increasingly present in modern industry and also in everyday life. Their applications range from health-related situations, for assistance to elderly people or in surgic…
APNet: Urban-level Scene Segmentation of Aerial Images and Point Clouds
Weijie Wei, Martin R. Oswald, Fatemeh Karimi Nejadasl +1
In this paper, we focus on semantic segmentation method for point clouds of urban scenes. Our fundamental concept revolves around the collaborative utilization of diverse scene rep…
SceneSplat++: A Large Dataset and Comprehensive Benchmark for Language Gaussian Splatting
Mengjiao Ma, Qi Ma, Yue Li +10
3D Gaussian Splatting (3DGS) serves as a highly performant and efficient encoding of scene geometry, appearance, and semantics. Moreover, grounding language in 3D scenes has proven…
ShadingNet: Image Intrinsics by Fine-Grained Shading Decomposition
Anil S. Baslamisli, Partha Das, Hoang-An Le +2
In general, intrinsic image decomposition algorithms interpret shading as one unified component including all photometric effects. As shading transitions are generally smoother tha…
Detect2Rank : Combining Object Detectors Using Learning to Rank
Sezer Karaoglu, Yang Liu, Theo Gevers
Object detection is an important research area in the field of computer vision. Many detection algorithms have been proposed. However, each object detector relies on specific assum…
FewViewGS: Gaussian Splatting with Few View Matching and Multi-stage Training
Ruihong Yin, Vladimir Yugay, Yue Li +2
The field of novel view synthesis from images has seen rapid advancements with the introduction of Neural Radiance Fields (NeRF) and more recently with 3D Gaussian Splatting. Gauss…
3DContextNet: K-d Tree Guided Hierarchical Learning of Point Clouds Using Local and Global Contextual Cues
Wei Zeng, Theo Gevers
Classification and segmentation of 3D point clouds are important tasks in computer vision. Because of the irregular nature of point clouds, most of the existing methods convert poi…
Geometry-guided Feature Learning and Fusion for Indoor Scene Reconstruction
Ruihong Yin, Sezer Karaoglu, Theo Gevers
In addition to color and textural information, geometry provides important cues for 3D scene reconstruction. However, current reconstruction methods only include geometry at the fe…
SceneTeller: Language-to-3D Scene Generation
BaÅak Melis Ãcal, Maxim Tatarchenko, Sezer Karaoglu +1
Designing high-quality indoor 3D scenes is important in many practical applications, such as room planning or game development. Conventionally, this has been a time-consuming proce…
LumiNet: Latent Intrinsics Meets Diffusion Models for Indoor Scene Relighting
Xiaoyan Xing, Konrad Groh, Sezer Karaoglu +2
We introduce LumiNet, a novel architecture that leverages generative models and latent intrinsic representations for effective lighting transfer. Given a source image and a target…
SDFit: 3D Object Pose and Shape by Fitting a Morphable SDF to a Single Image
Dimitrije AntiÄ, Georgios Paschalidis, Shashank Tripathi +3
Recovering 3D object pose and shape from a single image is a challenging and ill-posed problem. This is due to strong (self-)occlusions, depth ambiguities, the vast intra- and inte…
Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding
Yue Li, Qi Ma, Runyi Yang +8
While 3DGS has emerged as a high-fidelity scene representation, encoding rich, general-purpose features directly from its primitives remains under-explored. We address this gap by…
Three for one and one for three: Flow, Segmentation, and Surface Normals
Hoang-An Le, Anil S. Baslamisli, Thomas Mensink +1
Optical flow, semantic segmentation, and surface normals represent different information modalities, yet together they bring better cues for scene understanding problems. In this p…
CNN based Learning using Reflection and Retinex Models for Intrinsic Image Decomposition
Anil S. Baslamisli, Hoang-An Le, Theo Gevers
Most of the traditional work on intrinsic image decomposition rely on deriving priors about scene characteristics. On the other hand, recent research use deep learning models as in…
Fast SceneScript: Fast and Accurate Language-Based 3D Scene Understanding via Multi-Token Prediction
Ruihong Yin, Xuepeng Shi, Oleksandr Bailo +2
Recent perception-generalist approaches based on language models have achieved state-of-the-art results across diverse tasks, including 3D scene layout estimation and 3D object det…
PIE-Net: Photometric Invariant Edge Guided Network for Intrinsic Image Decomposition
Partha Das, Sezer Karaoglu, Theo Gevers
Intrinsic image decomposition is the process of recovering the image formation components (reflectance and shading) from an image. Previous methods employ either explicit priors to…
Unblur-SLAM: Dense Neural SLAM for Blurry Inputs
Qi Zhang, Denis Rozumny, Francesco Girlanda +4
We propose Unblur-SLAM, a novel RGB SLAM pipeline for sharp 3D reconstruction from blurred image inputs. In contrast to previous work, our approach is able to handle different type…
Edge-Centric Relational Reasoning for 3D Scene Graph Prediction
Yanni Ma, Hao Liu, Yulan Guo +2
3D scene graph prediction aims to abstract complex 3D environments into structured graphs consisting of objects and their pairwise relationships. Existing approaches typically adop…
Relighting as a Probe of Visual Priors via Augmented Latent Intrinsics
Xiaoyan Xing, Xiao Zhang, Sezer Karaoglu +2
Image-to-image relighting requires representations that separate illumination from scene properties while preserving dense geometry, material, and photometric cues. We use this tas…
Novel View Synthesis from Single Images via Point Cloud Transformation
Hoang-An Le, Thomas Mensink, Partha Das +1
In this paper the argument is made that for true novel view synthesis of objects, where the object can be synthesized from any viewpoint, an explicit 3D shape representation isdesi…
3D-AVS: LiDAR-based 3D Auto-Vocabulary Segmentation
Weijie Wei, Osman Ãlger, Fatemeh Karimi Nejadasl +2
Open-Vocabulary Segmentation (OVS) methods offer promising capabilities in detecting unseen object categories, but the category must be known and needs to be provided by a human, e…
Modeling Weather Uncertainty for Multi-weather Co-Presence Estimation
Qi Bi, Shaodi You, Theo Gevers
Images from outdoor scenes may be taken under various weather conditions. It is well studied that weather impacts the performance of computer vision algorithms and needs to be hand…
Inferring Point Clouds from Single Monocular Images by Depth Intermediation
Wei Zeng, Sezer Karaoglu, Theo Gevers
In this paper, we propose a pipeline to generate 3D point cloud of an object from a single-view RGB image. Most previous work predict the 3D point coordinates from single RGB image…
On the Benefit of Adversarial Training for Monocular Depth Estimation
Rick Groenendijk, Sezer Karaoglu, Theo Gevers +1
In this paper we address the benefit of adding adversarial training to the task of monocular depth estimation. A model can be trained in a self-supervised setting on stereo pairs o…
FVO: Fast Visual Odometry with Transformers
Vlardimir Yugay, Duy-Kien Nguyen, Theo Gevers +2
Hybrid pipelines that combine deep learning with classical optimization have established themselves as the dominant approach to visual odometry (VO). By integrating neural network…
HaarNet: Large-scale Linear-Morphological Hybrid Network for RGB-D Semantic Segmentation
Rick Groenendijk, Leo Dorst, Theo Gevers
Signals from different modalities each have their own combination algebra which affects their sampling processing. RGB is mostly linear; depth is a geometric signal following the o…
RealDiff: Real-world 3D Shape Completion using Self-Supervised Diffusion Models
BaÅak Melis Ãcal, Maxim Tatarchenko, Sezer Karaoglu +1
Point cloud completion aims to recover the complete 3D shape of an object from partial observations. While approaches relying on synthetic shape priors achieved promising results i…
EDEN: Multimodal Synthetic Dataset of Enclosed GarDEN Scenes
Hoang-An Le, Thomas Mensink, Partha Das +2
Multimodal large-scale datasets for outdoor scenes are mostly designed for urban driving problems. The scenes are highly structured and semantically different from scenarios seen i…
SIGNet: Intrinsic Image Decomposition by a Semantic and Invariant Gradient Driven Network for Indoor Scenes
Partha Das, Sezer Karaoglu, Arjan Gijsenij +1
Intrinsic image decomposition (IID) is an under-constrained problem. Therefore, traditional approaches use hand crafted priors to constrain the problem. However, these constraints…
Invariant Descriptors for Intrinsic Reflectance Optimization
Anil S. Baslamisli, Theo Gevers
Intrinsic image decomposition aims to factorize an image into albedo (reflectance) and shading (illumination) sub-components. Being ill-posed and under-constrained, it is a very ch…
Relational Prior Knowledge Graphs for Detection and Instance Segmentation
Osman Ãlger, Yu Wang, Ysbrand Galama +3
Humans have a remarkable ability to perceive and reason about the world around them by understanding the relationships between objects. In this paper, we investigate the effectiven…
Generative Models for Multi-Illumination Color Constancy
Partha Das, Yang Liu, Sezer Karaoglu +1
In this paper, the aim is multi-illumination color constancy. However, most of the existing color constancy methods are designed for single light sources. Furthermore, datasets for…
Multi-Loss Weighting with Coefficient of Variations
Rick Groenendijk, Sezer Karaoglu, Theo Gevers +1
Many interesting tasks in machine learning and computer vision are learned by optimising an objective function defined as a weighted linear combination of multiple losses. The fina…
Learning Content-enhanced Mask Transformer for Domain Generalized Urban-Scene Segmentation
Qi Bi, Shaodi You, Theo Gevers
Domain-generalized urban-scene semantic segmentation (USSS) aims to learn generalized semantic predictions across diverse urban-scene styles. Unlike domain gap challenges, USSS is…
GaussianBlender: Instant Stylization of 3D Gaussians with Disentangled Latent Spaces
Melis Ocal, Xiaoyan Xing, Yue Li +3
3D stylization is central to game development, virtual reality, and digital arts, where the demand for diverse assets calls for scalable methods that support fast, high-fidelity ma…
Ray-Distance Volume Rendering for Neural Scene Reconstruction
Ruihong Yin, Yunlu Chen, Sezer Karaoglu +1
Existing methods in neural scene reconstruction utilize the Signed Distance Function (SDF) to model the density function. However, in indoor scenes, the density computed from the S…
Improving Face Detection Performance with 3D-Rendered Synthetic Data
Jian Han, Sezer Karaoglu, Hoang-An Le +1
In this paper, we provide a synthetic data generator methodology with fully controlled, multifaceted variations based on a new 3D face dataset (3DU-Face). We customized synthetic d…
Retinex-Diffusion: On Controlling Illumination Conditions in Diffusion Models via Retinex Theory
Xiaoyan Xing, Vincent Tao Hu, Jan Hendrik Metzen +3
This paper introduces a novel approach to illumination manipulation in diffusion models, addressing the gap in conditional image generation with a focus on lighting conditions. We…
Intrinsic Image Decomposition Using Point Cloud Representation
Xiaoyan Xing, Konrad Groh, Sezer Karaoglu +1
The purpose of intrinsic decomposition is to separate an image into its albedo (reflective properties) and shading components (illumination properties). This is challenging because…