Publications (197)
Backpropagation-Friendly Eigendecomposition
Wei Wang, Zheng Dang, Yinlin Hu +2
Eigendecomposition (ED) is widely used in deep networks. However, the backpropagation of its results tends to be numerically unstable, whether using ED directly or approximating it…
OpenMaterial: A Large-scale Dataset of Complex Materials for 3D Reconstruction
Zheng Dang, Jialu Huang, Fei Wang +1
Recent advances in deep learning, such as neural radiance fields and implicit neural representations, have significantly advanced 3D reconstruction. However, accurately reconstruct…
Towards Robust Fine-grained Recognition by Maximal Separation of Discriminative Features
Krishna Kanth Nakka, Mathieu Salzmann
Adversarial attacks have been widely studied for general classification tasks, but remain unexplored in the context of fine-grained recognition, where the inter-class similarities…
OMH: Structured Sparsity via Optimally Matched Hierarchy for Unsupervised Semantic Segmentation
Baran Ozaydin, Tong Zhang, Deblina Bhattacharjee +2
Unsupervised Semantic Segmentation (USS) involves segmenting images without relying on predefined labels, aiming to alleviate the burden of extensive human labeling. Existing metho…
SegmentMeIfYouCan: A Benchmark for Anomaly Segmentation
Robin Chan, Krzysztof Lis, Svenja Uhlemeyer +6
State-of-the-art semantic or instance segmentation deep neural networks (DNNs) are usually trained on a closed set of semantic classes. As such, they are ill-equipped to handle pre…
Built-in Foreground/Background Prior for Weakly-Supervised Semantic Segmentation
Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Mathieu Salzmann +3
Pixel-level annotations are expensive and time consuming to obtain. Hence, weak supervision using only image tags could have a significant impact in semantic segmentation. Recently…
Beyond Gauss: Image-Set Matching on the Riemannian Manifold of PDFs
Mehrtash Harandi, Mathieu Salzmann, Mahsa Baktashmotlagh
State-of-the-art image-set matching techniques typically implicitly model each image-set with a Gaussian distribution. Here, we propose to go beyond these representations and model…
Temporally-Consistent Surface Reconstruction using Metrically-Consistent Atlases
Jan Bednarik, Noam Aigerman, Vladimir G. Kim +4
We propose a method for unsupervised reconstruction of a temporally-consistent sequence of surfaces from a sequence of time-evolving point clouds. It yields dense and semantically…
Adaptive Multi-step Refinement Network for Robust Point Cloud Registration
Zhi Chen, Yufan Ren, Tong Zhang +4
Point Cloud Registration (PCR) estimates the relative rigid transformation between two point clouds of the same scene. Despite significant progress with learning-based approaches,…
PCLs: Geometry-aware Neural Reconstruction of 3D Pose with Perspective Crop Layers
Frank Yu, Mathieu Salzmann, Pascal Fua +1
Local processing is an essential feature of CNNs and other neural network architectures - it is one of the reasons why they work so well on images where relevant information is, to…
KroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion Transformers
Yann Bouquet, Alireza Khodamoradi, Kristof Denolf +1
Post-training quantization (PTQ) of diffusion transformers (DiTs) to W4A4 severely degrades output quality, because activations entering each linear layer contain outliers that 4-b…
Effective Use of Synthetic Data for Urban Scene Semantic Segmentation
Fatemeh Sadat Saleh, Mohammad Sadegh Aliakbarian, Mathieu Salzmann +2
Training a deep network to perform semantic segmentation requires large amounts of labeled data. To alleviate the manual effort of annotating real images, researchers have investig…
Incorporating Network Built-in Priors in Weakly-supervised Semantic Segmentation
Fatemeh Sadat Saleh, Mohammad Sadegh Aliakbarian, Mathieu Salzmann +3
Pixel-level annotations are expensive and time consuming to obtain. Hence, weak supervision using only image tags could have a significant impact in semantic segmentation. Recently…
Perspective Flow Aggregation for Data-Limited 6D Object Pose Estimation
Yinlin Hu, Pascal Fua, Mathieu Salzmann
Most recent 6D object pose estimation methods, including unsupervised ones, require many real training images. Unfortunately, for some applications, such as those in space or deep…
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control
Mariam Hassan, Sebastian Stapf, Ahmad Rahimi +17
We present GEM, a Generalizable Ego-vision Multimodal world model that predicts future frames using a reference frame, sparse features, human poses, and ego-trajectories. Hence, ou…
Compression-aware Training of Deep Networks
Jose M. Alvarez, Mathieu Salzmann
In recent years, great progress has been made in a variety of application domains thanks to the development of increasingly deeper neural networks. Unfortunately, the huge number o…
Deep Action- and Context-Aware Sequence Learning for Activity Recognition and Anticipation
Mohammad Sadegh Aliakbarian, Fatemehsadat Saleh, Basura Fernando +3
Action recognition and anticipation are key to the success of many computer vision applications. Existing methods can roughly be grouped into those that extract global, context-awa…
Detecting Road Obstacles by Erasing Them
Krzysztof Lis, Sina Honari, Pascal Fua +1
Vehicles can encounter a myriad of obstacles on the road, and it is impossible to record them all beforehand to train a detector. Instead, we select image patches and inpaint them…
A Framework for Shape Analysis via Hilbert Space Embedding
Sadeep Jayasumana, Mathieu Salzmann, Hongdong Li +1
We propose a framework for 2D shape analysis using positive definite kernels defined on Kendall's shape manifold. Different representations of 2D shapes are known to generate diffe…
Modular Quantization-Aware Training for 6D Object Pose Estimation
Saqib Javed, Chengkun Li, Andrew Price +2
Edge applications, such as collaborative robotics and spacecraft rendezvous, demand efficient 6D object pose estimation on resource-constrained embedded platforms. Existing 6D pose…
Adaptive Low-Rank Kernel Subspace Clustering
Pan Ji, Ian Reid, Ravi Garg +2
In this paper, we present a kernel subspace clustering method that can handle non-linear models. In contrast to recent kernel subspace clustering methods which use predefined kerne…
DAAIN: Detection of Anomalous and Adversarial Input using Normalizing Flows
Samuel von BauÃnern, Johannes Otterbach, Adrian Loy +2
Despite much recent work, detecting out-of-distribution (OOD) inputs and adversarial attacks (AA) for computer vision models remains a challenge. In this work, we introduce a novel…
Evaluating the Search Phase of Neural Architecture Search
Kaicheng Yu, Christian Sciuto, Martin Jaggi +2
Neural Architecture Search (NAS) aims to facilitate the design of deep networks for new tasks. Existing techniques rely on two stages: searching over the architecture space and val…
Weight Space Representation Learning via Neural Field Adaptation
Zhuoqian Yang, Mathieu Salzmann, Sabine Süsstrunk
We investigate the potential of weights to serve as effective representations, focusing on neural fields. Our key insight is that constraining the optimization space through a pre-…
Efficient Linear Programming for Dense CRFs
Thalaiyasingam Ajanthan, Alban Desmaison, Rudy Bunel +3
The fully connected conditional random field (CRF) with Gaussian pairwise potentials has proven popular and effective for multi-class semantic segmentation. While the energy of a d…
Temporal Representation Learning on Monocular Videos for 3D Human Pose Estimation
Sina Honari, Victor Constantin, Helge Rhodin +2
In this paper we propose an unsupervised feature extraction method to capture temporal information on monocular videos, where we detect and encode subject of interest in each frame…
AttEntropy: On the Generalization Ability of Supervised Semantic Segmentation Transformers to New Objects in New Domains
Krzysztof Lis, Matthias Rottmann, Annika Mütze +3
In addition to impressive performance, vision transformers have demonstrated remarkable abilities to encode information they were not trained to extract. For example, this informat…
Learning to Fuse 2D and 3D Image Cues for Monocular Body Pose Estimation
Bugra Tekin, Pablo Márquez-Neila, Mathieu Salzmann +1
Most recent approaches to monocular 3D human pose estimation rely on Deep Learning. They typically involve regressing from an image to either 3D joint coordinates directly or 2D jo…
6Img-to-3D: Few-Image Large-Scale Outdoor Driving Scene Reconstruction
Théo Gieruc, Marius Kästingschäfer, Sebastian Bernhard +1
Current 3D reconstruction techniques struggle to infer unbounded scenes from a few images faithfully. Specifically, existing methods have high computational demands, require detail…
Vision Transformer Adapters for Generalizable Multitask Learning
Deblina Bhattacharjee, Sabine Süsstrunk, Mathieu Salzmann
We introduce the first multitasking vision transformer adapters that learn generalizable task affinities which can be applied to novel tasks and domains. Integrated into an off-the…
Kernel Coding: General Formulation and Special Cases
Mehrtash Harandi, Mathieu Salzmann
Representing images by compact codes has proven beneficial for many visual recognition tasks. Most existing techniques, however, perform this coding step directly in image feature…
Training Provably Robust Models by Polyhedral Envelope Regularization
Chen Liu, Mathieu Salzmann, Sabine Süsstrunk
Training certifiable neural networks enables one to obtain models with robustness guarantees against adversarial attacks. In this work, we introduce a framework to bound the advers…
Adversarial Parametric Pose Prior
Andrey Davydov, Anastasia Remizova, Victor Constantin +3
The Skinned Multi-Person Linear (SMPL) model can represent a human body by mapping pose and shape parameters to body meshes. This has been shown to facilitate inferring 3D human po…
What Stops Learning-based 3D Registration from Working in the Real World?
Zheng Dang, Lizhou Wang, Junning Qiu +2
Much progress has been made on the task of learning-based 3D point cloud registration, with existing methods yielding outstanding results on standard benchmarks, such as ModelNet40…
Fusing Local Similarities for Retrieval-based 3D Orientation Estimation of Unseen Objects
Chen Zhao, Yinlin Hu, Mathieu Salzmann
In this paper, we tackle the task of estimating the 3D orientation of previously-unseen objects from monocular images. This task contrasts with the one considered by most existing…
3D Registration for Self-Occluded Objects in Context
Zheng Dang, Fei Wang, Mathieu Salzmann
While much progress has been made on the task of 3D point cloud registration, there still exists no learning-based method able to estimate the 6D pose of an object observed by a 2.…
Imposing Hard Constraints on Deep Networks: Promises and Limitations
Pablo Márquez-Neila, Mathieu Salzmann, Pascal Fua
Imposing constraints on the output of a Deep Neural Net is one way to improve the quality of its predictions while loosening the requirements for labeled training data. Such constr…
Source-Free Domain-Invariant Performance Prediction
Ekaterina Khramtsova, Mahsa Baktashmotlagh, Guido Zuccon +2
Accurately estimating model performance poses a significant challenge, particularly in scenarios where the source and target domains follow different data distributions. Most exist…
From Manifold to Manifold: Geometry-Aware Dimensionality Reduction for SPD Matrices
Mehrtash T. Harandi, Mathieu Salzmann, Richard Hartley
Representing images and videos with Symmetric Positive Definite (SPD) matrices and considering the Riemannian geometry of the resulting space has proven beneficial for many recogni…
Calibration-Reasoning Framework for Descriptive Speech Quality Assessment
Elizaveta Kostenok, Mathieu Salzmann, Milos Cernak
Explainable speech quality assessment requires moving beyond Mean Opinion Scores (MOS) to analyze underlying perceptual dimensions. To address this, we introduce a novel post-train…
Using Motion Cues to Supervise Single-Frame Body Pose and Shape Estimation in Low Data Regimes
Andrey Davydov, Alexey Sidnev, Artsiom Sanakoyeu +3
When enough annotated training data is available, supervised deep-learning algorithms excel at estimating human body pose and shape using a single camera. The effects of too little…
Progressive Correspondence Pruning by Consensus Learning
Chen Zhao, Yixiao Ge, Feng Zhu +3
Correspondence selection aims to correctly select the consistent matches (inliers) from an initial set of putative correspondences. The selection is challenging since putative matc…
Bregman Divergences for Infinite Dimensional Covariance Matrices
Mehrtash Harandi, Mathieu Salzmann, Fatih Porikli
We introduce an approach to computing and comparing Covariance Descriptors (CovDs) in infinite-dimensional spaces. CovDs have become increasingly popular to address classification…
Estimating People Flows to Better Count Them in Crowded Scenes
Weizhe Liu, Mathieu Salzmann, Pascal Fua
Modern methods for counting people in crowded scenes rely on deep networks to estimate people densities in individual images. As such, only very few take advantage of temporal cons…
Dimensionality Reduction on SPD Manifolds: The Emergence of Geometry-Aware Methods
Mehrtash Harandi, Mathieu Salzmann, Richard Hartley
Representing images and videos with Symmetric Positive Definite (SPD) matrices, and considering the Riemannian geometry of the resulting space, has been shown to yield high discrim…
How to Train Your Super-Net: An Analysis of Training Heuristics in Weight-Sharing NAS
Kaicheng Yu, Rene Ranftl, Mathieu Salzmann
Weight sharing promises to make neural architecture search (NAS) tractable even on commodity hardware. Existing methods in this space rely on a diverse set of heuristics to design…
Demystifying Singular Defects in Large Language Models
Haoqi Wang, Tong Zhang, Mathieu Salzmann
Large transformer models are known to produce high-norm tokens. In vision transformers (ViTs), such tokens have been mathematically modeled through the singular vectors of the line…
Learning Factorized Representations for Open-set Domain Adaptation
Mahsa Baktashmotlagh, Masoud Faraki, Tom Drummond +1
Domain adaptation for visual recognition has undergone great progress in the past few years. Nevertheless, most existing methods work in the so-called closed-set scenario, assuming…
Unsupervised 3D Keypoint Discovery with Multi-View Geometry
Sina Honari, Chen Zhao, Mathieu Salzmann +1
Analyzing and training 3D body posture models depend heavily on the availability of joint labels that are commonly acquired through laborious manual annotation of body joints or vi…
Deep Subspace Clustering Networks
Pan Ji, Tong Zhang, Hongdong Li +2
We present a novel deep neural network architecture for unsupervised subspace clustering. This architecture is built upon deep auto-encoders, which non-linearly map the input data…
MulT: An End-to-End Multitask Learning Transformer
Deblina Bhattacharjee, Tong Zhang, Sabine Süsstrunk +1
We propose an end-to-end Multitask Learning Transformer framework, named MulT, to simultaneously learn multiple high-level vision tasks, including depth estimation, semantic segmen…
Indirect Local Attacks for Context-aware Semantic Segmentation Networks
Krishna Kanth Nakka, Mathieu Salzmann
Recently, deep networks have achieved impressive semantic segmentation performance, in particular thanks to their use of larger contextual information. In this paper, we show that…
Knowledge Distillation for 6D Pose Estimation by Aligning Distributions of Local Predictions
Shuxuan Guo, Yinlin Hu, Jose M. Alvarez +1
Knowledge distillation facilitates the training of a compact student network by using a deep teacher one. While this has achieved great success in many tasks, it remains completely…
Wide-Depth-Range 6D Object Pose Estimation in Space
Yinlin Hu, Sebastien Speierer, Wenzel Jakob +2
6D pose estimation in space poses unique challenges that are not commonly encountered in the terrestrial setting. One of the most striking differences is the lack of atmospheric sc…
3D Pose Based Feedback for Physical Exercises
Ziyi Zhao, Sena Kiciroglu, Hugues Vinzant +4
Unsupervised self-rehabilitation exercises and physical training can cause serious injuries if performed incorrectly. We introduce a learning-based framework that identifies the mi…
Overcoming Multi-Model Forgetting
Yassine Benyahia, Kaicheng Yu, Kamil Bennani-Smires +4
We identify a phenomenon, which we refer to as multi-model forgetting, that occurs when sequentially training multiple deep networks with partially-shared parameters; the performan…
Robust Differentiable SVD
Wei Wang, Zheng Dang, Yinlin Hu +2
Eigendecomposition of symmetric matrices is at the heart of many computer vision algorithms. However, the derivatives of the eigenvectors tend to be numerically unstable, whether u…
On the Impact of Hard Adversarial Instances on Overfitting in Adversarial Training
Chen Liu, Zhichao Huang, Mathieu Salzmann +2
Adversarial training is a popular method to robustify models against adversarial attacks. However, it exhibits much more severe overfitting than training on clean inputs. In this w…
Learning 3D-3D Correspondences for One-shot Partial-to-partial Registration
Zheng Dang, Fei Wang, Mathieu Salzmann
While 3D-3D registration is traditionally tacked by optimization-based methods, recent work has shown that learning-based techniques could achieve faster and more robust results. I…
Single-Stage 6D Object Pose Estimation
Yinlin Hu, Pascal Fua, Wei Wang +1
Most recent 6D pose estimation frameworks first rely on a deep network to establish correspondences between 3D object keypoints and 2D image locations and then use a variant of a R…
Unlocking Comics: The AI4VA Dataset for Visual Understanding
Peter Grönquist, Deblina Bhattacharjee, Bahar Aydemir +4
In the evolving landscape of deep learning, there is a pressing need for more comprehensive datasets capable of training models across multiple modalities. Concurrently, in digital…
Contact-aware Human Motion Forecasting
Wei Mao, Miaomiao Liu, Richard Hartley +1
In this paper, we tackle the task of scene-aware 3D human motion forecasting, which consists of predicting future human poses given a 3D scene and a past human motion. A key challe…
3D Single-object Tracking in Point Clouds with High Temporal Variation
Qiao Wu, Kun Sun, Pei An +3
The high temporal variation of the point clouds is the key challenge of 3D single-object tracking (3D SOT). Existing approaches rely on the assumption that the shape variation of t…
Linear-Covariance Loss for End-to-End Learning of 6D Pose Estimation
Fulin Liu, Yinlin Hu, Mathieu Salzmann
Most modern image-based 6D object pose estimation methods learn to predict 2D-3D correspondences, from which the pose can be obtained using a PnP solver. Because of the non-differe…
Long Term Motion Prediction Using Keyposes
Sena Kiciroglu, Wei Wang, Mathieu Salzmann +1
Long term human motion prediction is essential in safety-critical applications such as human-robot interaction and autonomous driving. In this paper we show that to achieve long te…
Coherent and Multi-modality Image Inpainting via Latent Space Optimization
Lingzhi Pan, Tong Zhang, Bingyuan Chen +4
With the advancements in denoising diffusion probabilistic models (DDPMs), image inpainting has significantly evolved from merely filling information based on nearby regions to gen…
Robust RGB-based 6-DoF Pose Estimation without Real Pose Annotations
Zhigang Li, Yinlin Hu, Mathieu Salzmann +1
While much progress has been made in 6-DoF object pose estimation from a single RGB image, the current leading approaches heavily rely on real-annotation data. As such, they remain…
FastPose-ViT: A Vision Transformer for Real-Time Spacecraft Pose Estimation
Pierre Ancey, Andrew Price, Saqib Javed +1
Estimating the 6-degrees-of-freedom (6DoF) pose of a spacecraft from a single image is critical for autonomous operations like in-orbit servicing and space debris removal. Existing…
Attention-based Domain Adaptation for Single Stage Detectors
Vidit Vidit, Mathieu Salzmann
While domain adaptation has been used to improve the performance of object detectors when the training and test data follow different distributions, previous work has mostly focuse…
Learning Variations in Human Motion via Mix-and-Match Perturbation
Mohammad Sadegh Aliakbarian, Fatemeh Sadat Saleh, Mathieu Salzmann +3
Human motion prediction is a stochastic process: Given an observed sequence of poses, multiple future motions are plausible. Existing approaches to modeling this stochasticity typi…
VIENA2: A Driving Anticipation Dataset
Mohammad Sadegh Aliakbarian, Fatemeh Sadat Saleh, Mathieu Salzmann +3
Action anticipation is critical in scenarios where one needs to react before the action is finalized. This is, for instance, the case in automated driving, where a car needs to, e.…
CLOAF: CoLlisiOn-Aware Human Flow
Andrey Davydov, Martin Engilberge, Mathieu Salzmann +1
Even the best current algorithms for estimating body 3D shape and pose yield results that include body self-intersections. In this paper, we present CLOAF, which exploits the diffe…
Q-Drift: Quantization-Aware Drift Correction for Diffusion Model Sampling
Sooyoung Ryu, Mathieu Salzmann, Saqib Javed
Post-training quantization (PTQ) is a practical path to deploy large diffusion models, but quantization noise can accumulate over the denoising trajectory and degrade generation qu…
Learning to Weight Parameters for Training Data Attribution
Shuangqi Li, Hieu Le, Jingyi Xu +1
We study gradient-based data attribution, aiming to identify which training examples most influence a given output. Existing methods for this task either treat network parameters u…
Sample and Filter: Nonparametric Scene Parsing via Efficient Filtering
Mohammad Najafi, Sarah Taghavi Namin, Mathieu Salzmann +1
Scene parsing has attracted a lot of attention in computer vision. While parametric models have proven effective for this task, they cannot easily incorporate new training data. By…
GarNet++: Improving Fast and Accurate Static3D Cloth Draping by Curvature Loss
Erhan Gundogdu, Victor Constantin, Shaifali Parashar +4
In this paper, we tackle the problem of static 3D cloth draping on virtual human bodies. We introduce a two-stream deep network model that produces a visually plausible draping of…
Statistically Motivated Second Order Pooling
Kaicheng Yu, Mathieu Salzmann
Second-order pooling, a.k.a.~bilinear pooling, has proven effective for deep learning based visual recognition. However, the resulting second-order networks yield a final represent…
Shape Interaction Matrix Revisited and Robustified: Efficient Subspace Clustering with Corrupted and Incomplete Data
Pan Ji, Mathieu Salzmann, Hongdong Li
The Shape Interaction Matrix (SIM) is one of the earliest approaches to performing subspace clustering (i.e., separating points drawn from a union of subspaces). In this paper, we…
Kernel Methods on Riemannian Manifolds with Gaussian RBF Kernels
Sadeep Jayasumana, Richard Hartley, Mathieu Salzmann +2
In this paper, we develop an approach to exploiting kernel methods with manifold-valued data. In many computer vision problems, the data can be naturally represented as points on a…
QT-DoG: Quantization-aware Training for Domain Generalization
Saqib Javed, Hieu Le, Mathieu Salzmann
A key challenge in Domain Generalization (DG) is preventing overfitting to source domains, which can be mitigated by finding flatter minima in the loss landscape. In this work, we…
Temporally-Transferable Perturbations: Efficient, One-Shot Adversarial Attacks for Online Visual Object Trackers
Krishna Kanth Nakka, Mathieu Salzmann
In recent years, the trackers based on Siamese networks have emerged as highly effective and efficient for visual object tracking (VOT). While these methods were shown to be vulner…
Generalize or Detect? Towards Robust Semantic Segmentation Under Multiple Distribution Shifts
Zhitong Gao, Bingnan Li, Mathieu Salzmann +1
In open-world scenarios, where both novel classes and domains may exist, an ideal segmentation model should detect anomaly classes for safety and generalize to new domains. However…
SINDER: Repairing the Singular Defects of DINOv2
Haoqi Wang, Tong Zhang, Mathieu Salzmann
Vision Transformer models trained on large-scale datasets, although effective, often exhibit artifacts in the patch token they extract. While such defects can be alleviated by re-t…
3D-Aware Hypothesis & Verification for Generalizable Relative Object Pose Estimation
Chen Zhao, Tong Zhang, Mathieu Salzmann
Prior methods that tackle the problem of generalizable object pose estimation highly rely on having dense views of the unseen object. By contrast, we address the scenario where onl…
Controlling the Fidelity and Diversity of Deep Generative Models via Pseudo Density
Shuangqi Li, Chen Liu, Tong Zhang +3
We introduce an approach to bias deep generative models, such as GANs and diffusion models, towards generating data with either enhanced fidelity or increased diversity. Our approa…
Leverage Your Local and Global Representations: A New Self-Supervised Learning Strategy
Tong Zhang, Congpei Qiu, Wei Ke +2
Self-supervised learning (SSL) methods aim to learn view-invariant representations by maximizing the similarity between the features extracted from different crops of the same imag…
Beyond One Glance: Gated Recurrent Architecture for Hand Segmentation
Wei Wang, Kaicheng Yu, Joachim Hugonot +2
As mixed reality is gaining increased momentum, the development of effective and efficient solutions to egocentric hand segmentation is becoming critical. Traditional segmentation…
Soft Correspondences in Multimodal Scene Parsing
Sarah Taghavi Namin, Mohammad Najafi, Mathieu Salzmann +1
Exploiting multiple modalities for semantic scene parsing has been shown to improve accuracy over the singlemodality scenario. However multimodal datasets often suffer from problem…
On the Loss Landscape of Adversarial Training: Identifying Challenges and How to Overcome Them
Chen Liu, Mathieu Salzmann, Tao Lin +2
We analyze the influence of adversarial training on the loss landscape of machine learning models. To this end, we first provide analytical studies of the properties of adversarial…
Semantic-Aware Depth Super-Resolution in Outdoor Scenes
Miaomiao Liu, Mathieu Salzmann, Xuming He
While depth sensors are becoming increasingly popular, their spatial resolution often remains limited. Depth super-resolution therefore emerged as a solution to this problem. Despi…
When VLAD met Hilbert
Mehrtash Harandi, Mathieu Salzmann, Fatih Porikli
Vectors of Locally Aggregated Descriptors (VLAD) have emerged as powerful image/video representations that compete with or even outperform state-of-the-art approaches on many chall…
Dyadic Human Motion Prediction
Isinsu Katircioglu, Costa Georgantas, Mathieu Salzmann +1
Prior work on human motion forecasting has mostly focused on predicting the future motion of single subjects in isolation from their past pose sequence. In the presence of closely…
TempSAL -- Uncovering Temporal Information for Deep Saliency Prediction
Bahar Aydemir, Ludo Hoffstetter, Tong Zhang +2
Deep saliency prediction algorithms complement the object recognition features, they typically rely on additional information, such as scene context, semantic relationships, gaze d…
Towards Self-Supervised Covariance Estimation in Deep Heteroscedastic Regression
Megh Shukla, Aziz Shameem, Mathieu Salzmann +1
Deep heteroscedastic regression models the mean and covariance of the target distribution through neural networks. The challenge arises from heteroscedasticity, which implies that…
Contextually Plausible and Diverse 3D Human Motion Prediction
Sadegh Aliakbarian, Fatemeh Sadat Saleh, Lars Petersson +2
We tackle the task of diverse 3D human motion prediction, that is, forecasting multiple plausible future 3D poses given a sequence of observed 3D poses. In this context, a popular…
Iteratively Reweighted Graph Cut for Multi-label MRFs with Non-convex Priors
Thalaiyasingam Ajanthan, Richard Hartley, Mathieu Salzmann +1
While widely acknowledged as highly effective in computer vision, multi-label MRFs with non-convex priors are difficult to optimize. To tackle this, we introduce an algorithm that…
LoRaQ: Optimized Low Rank Approximation for 4-bit Quantization
Yann Bouquet, Alireza Khodamoradi, Sophie Yáng Shen +2
Post-training quantization (PTQ) is essential for deploying large diffusion transformers on resource-constrained hardware, but aggressive 4-bit quantization significantly degrades…
Neural Scene Decomposition for Multi-Person Motion Capture
Helge Rhodin, Victor Constantin, Isinsu Katircioglu +2
Learning general image representations has proven key to the success of many computer vision tasks. For example, many approaches to image understanding problems rely on deep networ…
Recurrent U-Net for Resource-Constrained Segmentation
Wei Wang, Kaicheng Yu, Joachim Hugonot +2
State-of-the-art segmentation methods rely on very deep networks that are not always easy to train without very large training datasets and tend to be relatively slow to run on sta…
Deep learning to improve the discovery of near-Earth asteroids in the Zwicky Transient Facility
Belén Yu Irureta-Goyena, George Helou, Jean-Paul Kneib +7
We present a novel pipeline that uses a convolutional neural network (CNN) to improve the detection capability of near-Earth asteroids (NEAs) in the context of planetary defense. O…