papers

Publications (197)

cs.LG2019

Backpropagation-Friendly Eigendecomposition

Wei Wang, Zheng Dang, Yinlin Hu +2

Eigendecomposition (ED) is widely used in deep networks. However, the backpropagation of its results tends to be numerically unstable, whether using ED directly or approximating it…

cs.CV2025

OpenMaterial: A Large-scale Dataset of Complex Materials for 3D Reconstruction

Zheng Dang, Jialu Huang, Fei Wang +1

Recent advances in deep learning, such as neural radiance fields and implicit neural representations, have significantly advanced 3D reconstruction. However, accurately reconstruct…

cs.CV2020

Towards Robust Fine-grained Recognition by Maximal Separation of Discriminative Features

Krishna Kanth Nakka, Mathieu Salzmann

Adversarial attacks have been widely studied for general classification tasks, but remain unexplored in the context of fine-grained recognition, where the inter-class similarities…

cs.CV2024

OMH: Structured Sparsity via Optimally Matched Hierarchy for Unsupervised Semantic Segmentation

Baran Ozaydin, Tong Zhang, Deblina Bhattacharjee +2

Unsupervised Semantic Segmentation (USS) involves segmenting images without relying on predefined labels, aiming to alleviate the burden of extensive human labeling. Existing metho…

cs.CV2021

SegmentMeIfYouCan: A Benchmark for Anomaly Segmentation

Robin Chan, Krzysztof Lis, Svenja Uhlemeyer +6

State-of-the-art semantic or instance segmentation deep neural networks (DNNs) are usually trained on a closed set of semantic classes. As such, they are ill-equipped to handle pre…

cs.CV2016

Built-in Foreground/Background Prior for Weakly-Supervised Semantic Segmentation

Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Mathieu Salzmann +3

Pixel-level annotations are expensive and time consuming to obtain. Hence, weak supervision using only image tags could have a significant impact in semantic segmentation. Recently…

cs.CV2015

Beyond Gauss: Image-Set Matching on the Riemannian Manifold of PDFs

Mehrtash Harandi, Mathieu Salzmann, Mahsa Baktashmotlagh

State-of-the-art image-set matching techniques typically implicitly model each image-set with a Gaussian distribution. Here, we propose to go beyond these representations and model…

cs.CV2021

Temporally-Consistent Surface Reconstruction using Metrically-Consistent Atlases

Jan Bednarik, Noam Aigerman, Vladimir G. Kim +4

We propose a method for unsupervised reconstruction of a temporally-consistent sequence of surfaces from a sequence of time-evolving point clouds. It yields dense and semantically…

cs.CV2025

Adaptive Multi-step Refinement Network for Robust Point Cloud Registration

Zhi Chen, Yufan Ren, Tong Zhang +4

Point Cloud Registration (PCR) estimates the relative rigid transformation between two point clouds of the same scene. Despite significant progress with learning-based approaches,…

cs.CV2021

PCLs: Geometry-aware Neural Reconstruction of 3D Pose with Perspective Crop Layers

Frank Yu, Mathieu Salzmann, Pascal Fua +1

Local processing is an essential feature of CNNs and other neural network architectures - it is one of the reasons why they work so well on images where relevant information is, to…

cs.LG2026

KroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion Transformers

Yann Bouquet, Alireza Khodamoradi, Kristof Denolf +1

Post-training quantization (PTQ) of diffusion transformers (DiTs) to W4A4 severely degrades output quality, because activations entering each linear layer contain outliers that 4-b…

cs.CV2018

Effective Use of Synthetic Data for Urban Scene Semantic Segmentation

Fatemeh Sadat Saleh, Mohammad Sadegh Aliakbarian, Mathieu Salzmann +2

Training a deep network to perform semantic segmentation requires large amounts of labeled data. To alleviate the manual effort of annotating real images, researchers have investig…

cs.CV2017

Incorporating Network Built-in Priors in Weakly-supervised Semantic Segmentation

Fatemeh Sadat Saleh, Mohammad Sadegh Aliakbarian, Mathieu Salzmann +3

Pixel-level annotations are expensive and time consuming to obtain. Hence, weak supervision using only image tags could have a significant impact in semantic segmentation. Recently…

cs.CV2022

Perspective Flow Aggregation for Data-Limited 6D Object Pose Estimation

Yinlin Hu, Pascal Fua, Mathieu Salzmann

Most recent 6D object pose estimation methods, including unsupervised ones, require many real training images. Unfortunately, for some applications, such as those in space or deep…

cs.CV2024

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

Mariam Hassan, Sebastian Stapf, Ahmad Rahimi +17

We present GEM, a Generalizable Ego-vision Multimodal world model that predicts future frames using a reference frame, sparse features, human poses, and ego-trajectories. Hence, ou…

cs.CV2017

Compression-aware Training of Deep Networks

Jose M. Alvarez, Mathieu Salzmann

In recent years, great progress has been made in a variety of application domains thanks to the development of increasingly deeper neural networks. Unfortunately, the huge number o…

cs.CV2016

Deep Action- and Context-Aware Sequence Learning for Activity Recognition and Anticipation

Mohammad Sadegh Aliakbarian, Fatemehsadat Saleh, Basura Fernando +3

Action recognition and anticipation are key to the success of many computer vision applications. Existing methods can roughly be grouped into those that extract global, context-awa…

cs.CV2023

Detecting Road Obstacles by Erasing Them

Krzysztof Lis, Sina Honari, Pascal Fua +1

Vehicles can encounter a myriad of obstacles on the road, and it is impossible to record them all beforehand to train a detector. Instead, we select image patches and inpaint them…

cs.CV2014

A Framework for Shape Analysis via Hilbert Space Embedding

Sadeep Jayasumana, Mathieu Salzmann, Hongdong Li +1

We propose a framework for 2D shape analysis using positive definite kernels defined on Kendall's shape manifold. Different representations of 2D shapes are known to generate diffe…

cs.CV2024

Modular Quantization-Aware Training for 6D Object Pose Estimation

Saqib Javed, Chengkun Li, Andrew Price +2

Edge applications, such as collaborative robotics and spacecraft rendezvous, demand efficient 6D object pose estimation on resource-constrained embedded platforms. Existing 6D pose…

cs.CV2019

Adaptive Low-Rank Kernel Subspace Clustering

Pan Ji, Ian Reid, Ravi Garg +2

In this paper, we present a kernel subspace clustering method that can handle non-linear models. In contrast to recent kernel subspace clustering methods which use predefined kerne…

cs.CV2021

DAAIN: Detection of Anomalous and Adversarial Input using Normalizing Flows

Samuel von Baußnern, Johannes Otterbach, Adrian Loy +2

Despite much recent work, detecting out-of-distribution (OOD) inputs and adversarial attacks (AA) for computer vision models remains a challenge. In this work, we introduce a novel…

cs.LG2019

Evaluating the Search Phase of Neural Architecture Search

Kaicheng Yu, Christian Sciuto, Martin Jaggi +2

Neural Architecture Search (NAS) aims to facilitate the design of deep networks for new tasks. Existing techniques rely on two stages: searching over the architecture space and val…

cs.LG2026

Weight Space Representation Learning via Neural Field Adaptation

Zhuoqian Yang, Mathieu Salzmann, Sabine Süsstrunk

We investigate the potential of weights to serve as effective representations, focusing on neural fields. Our key insight is that constraining the optimization space through a pre-…

cs.CV2017

Efficient Linear Programming for Dense CRFs

Thalaiyasingam Ajanthan, Alban Desmaison, Rudy Bunel +3

The fully connected conditional random field (CRF) with Gaussian pairwise potentials has proven popular and effective for multi-class semantic segmentation. While the energy of a d…

cs.CV2022

Temporal Representation Learning on Monocular Videos for 3D Human Pose Estimation

Sina Honari, Victor Constantin, Helge Rhodin +2

In this paper we propose an unsupervised feature extraction method to capture temporal information on monocular videos, where we detect and encode subject of interest in each frame…

cs.CV2024

AttEntropy: On the Generalization Ability of Supervised Semantic Segmentation Transformers to New Objects in New Domains

Krzysztof Lis, Matthias Rottmann, Annika Mütze +3

In addition to impressive performance, vision transformers have demonstrated remarkable abilities to encode information they were not trained to extract. For example, this informat…

cs.CV2017

Learning to Fuse 2D and 3D Image Cues for Monocular Body Pose Estimation

Bugra Tekin, Pablo Márquez-Neila, Mathieu Salzmann +1

Most recent approaches to monocular 3D human pose estimation rely on Deep Learning. They typically involve regressing from an image to either 3D joint coordinates directly or 2D jo…

cs.CV2025

6Img-to-3D: Few-Image Large-Scale Outdoor Driving Scene Reconstruction

Théo Gieruc, Marius Kästingschäfer, Sebastian Bernhard +1

Current 3D reconstruction techniques struggle to infer unbounded scenes from a few images faithfully. Specifically, existing methods have high computational demands, require detail…

cs.CV2023

Vision Transformer Adapters for Generalizable Multitask Learning

Deblina Bhattacharjee, Sabine Süsstrunk, Mathieu Salzmann

We introduce the first multitasking vision transformer adapters that learn generalizable task affinities which can be applied to novel tasks and domains. Integrated into an off-the…

cs.CV2014

Kernel Coding: General Formulation and Special Cases

Mehrtash Harandi, Mathieu Salzmann

Representing images by compact codes has proven beneficial for many visual recognition tasks. Most existing techniques, however, perform this coding step directly in image feature…

cs.LG2021

Training Provably Robust Models by Polyhedral Envelope Regularization

Chen Liu, Mathieu Salzmann, Sabine Süsstrunk

Training certifiable neural networks enables one to obtain models with robustness guarantees against adversarial attacks. In this work, we introduce a framework to bound the advers…

cs.CV2021

Adversarial Parametric Pose Prior

Andrey Davydov, Anastasia Remizova, Victor Constantin +3

The Skinned Multi-Person Linear (SMPL) model can represent a human body by mapping pose and shape parameters to body meshes. This has been shown to facilitate inferring 3D human po…

cs.CV2021

What Stops Learning-based 3D Registration from Working in the Real World?

Zheng Dang, Lizhou Wang, Junning Qiu +2

Much progress has been made on the task of learning-based 3D point cloud registration, with existing methods yielding outstanding results on standard benchmarks, such as ModelNet40…

cs.CV2022

Fusing Local Similarities for Retrieval-based 3D Orientation Estimation of Unseen Objects

Chen Zhao, Yinlin Hu, Mathieu Salzmann

In this paper, we tackle the task of estimating the 3D orientation of previously-unseen objects from monocular images. This task contrasts with the one considered by most existing…

cs.CV2020

3D Registration for Self-Occluded Objects in Context

Zheng Dang, Fei Wang, Mathieu Salzmann

While much progress has been made on the task of 3D point cloud registration, there still exists no learning-based method able to estimate the 6D pose of an object observed by a 2.…

cs.CV2017

Imposing Hard Constraints on Deep Networks: Promises and Limitations

Pablo Márquez-Neila, Mathieu Salzmann, Pascal Fua

Imposing constraints on the output of a Deep Neural Net is one way to improve the quality of its predictions while loosening the requirements for labeled training data. Such constr…

cs.CV2024

Source-Free Domain-Invariant Performance Prediction

Ekaterina Khramtsova, Mahsa Baktashmotlagh, Guido Zuccon +2

Accurately estimating model performance poses a significant challenge, particularly in scenarios where the source and target domains follow different data distributions. Most exist…

cs.CV2014

From Manifold to Manifold: Geometry-Aware Dimensionality Reduction for SPD Matrices

Mehrtash T. Harandi, Mathieu Salzmann, Richard Hartley

Representing images and videos with Symmetric Positive Definite (SPD) matrices and considering the Riemannian geometry of the resulting space has proven beneficial for many recogni…

eess.AS2026

Calibration-Reasoning Framework for Descriptive Speech Quality Assessment

Elizaveta Kostenok, Mathieu Salzmann, Milos Cernak

Explainable speech quality assessment requires moving beyond Mean Opinion Scores (MOS) to analyze underlying perceptual dimensions. To address this, we introduce a novel post-train…

cs.CV2024

Using Motion Cues to Supervise Single-Frame Body Pose and Shape Estimation in Low Data Regimes

Andrey Davydov, Alexey Sidnev, Artsiom Sanakoyeu +3

When enough annotated training data is available, supervised deep-learning algorithms excel at estimating human body pose and shape using a single camera. The effects of too little…

cs.CV2021

Progressive Correspondence Pruning by Consensus Learning

Chen Zhao, Yixiao Ge, Feng Zhu +3

Correspondence selection aims to correctly select the consistent matches (inliers) from an initial set of putative correspondences. The selection is challenging since putative matc…

cs.CV2014

Bregman Divergences for Infinite Dimensional Covariance Matrices

Mehrtash Harandi, Mathieu Salzmann, Fatih Porikli

We introduce an approach to computing and comparing Covariance Descriptors (CovDs) in infinite-dimensional spaces. CovDs have become increasingly popular to address classification…

cs.CV2020

Estimating People Flows to Better Count Them in Crowded Scenes

Weizhe Liu, Mathieu Salzmann, Pascal Fua

Modern methods for counting people in crowded scenes rely on deep networks to estimate people densities in individual images. As such, only very few take advantage of temporal cons…

cs.CV2016

Dimensionality Reduction on SPD Manifolds: The Emergence of Geometry-Aware Methods

Mehrtash Harandi, Mathieu Salzmann, Richard Hartley

Representing images and videos with Symmetric Positive Definite (SPD) matrices, and considering the Riemannian geometry of the resulting space, has been shown to yield high discrim…

cs.LG2020

How to Train Your Super-Net: An Analysis of Training Heuristics in Weight-Sharing NAS

Kaicheng Yu, Rene Ranftl, Mathieu Salzmann

Weight sharing promises to make neural architecture search (NAS) tractable even on commodity hardware. Existing methods in this space rely on a diverse set of heuristics to design…

cs.CL2025

Demystifying Singular Defects in Large Language Models

Haoqi Wang, Tong Zhang, Mathieu Salzmann

Large transformer models are known to produce high-norm tokens. In vision transformers (ViTs), such tokens have been mathematically modeled through the singular vectors of the line…

cs.CV2018

Learning Factorized Representations for Open-set Domain Adaptation

Mahsa Baktashmotlagh, Masoud Faraki, Tom Drummond +1

Domain adaptation for visual recognition has undergone great progress in the past few years. Nevertheless, most existing methods work in the so-called closed-set scenario, assuming…

cs.CV2024

Unsupervised 3D Keypoint Discovery with Multi-View Geometry

Sina Honari, Chen Zhao, Mathieu Salzmann +1

Analyzing and training 3D body posture models depend heavily on the availability of joint labels that are commonly acquired through laborious manual annotation of body joints or vi…

cs.CV2017

Deep Subspace Clustering Networks

Pan Ji, Tong Zhang, Hongdong Li +2

We present a novel deep neural network architecture for unsupervised subspace clustering. This architecture is built upon deep auto-encoders, which non-linearly map the input data…

cs.CV2022

MulT: An End-to-End Multitask Learning Transformer

Deblina Bhattacharjee, Tong Zhang, Sabine Süsstrunk +1

We propose an end-to-end Multitask Learning Transformer framework, named MulT, to simultaneously learn multiple high-level vision tasks, including depth estimation, semantic segmen…

cs.CV2019

Indirect Local Attacks for Context-aware Semantic Segmentation Networks

Krishna Kanth Nakka, Mathieu Salzmann

Recently, deep networks have achieved impressive semantic segmentation performance, in particular thanks to their use of larger contextual information. In this paper, we show that…

cs.CV2022

Knowledge Distillation for 6D Pose Estimation by Aligning Distributions of Local Predictions

Shuxuan Guo, Yinlin Hu, Jose M. Alvarez +1

Knowledge distillation facilitates the training of a compact student network by using a deep teacher one. While this has achieved great success in many tasks, it remains completely…

cs.CV2021

Wide-Depth-Range 6D Object Pose Estimation in Space

Yinlin Hu, Sebastien Speierer, Wenzel Jakob +2

6D pose estimation in space poses unique challenges that are not commonly encountered in the terrestrial setting. One of the most striking differences is the lack of atmospheric sc…

cs.CV2022

3D Pose Based Feedback for Physical Exercises

Ziyi Zhao, Sena Kiciroglu, Hugues Vinzant +4

Unsupervised self-rehabilitation exercises and physical training can cause serious injuries if performed incorrectly. We introduce a learning-based framework that identifies the mi…

cs.LG2019

Overcoming Multi-Model Forgetting

Yassine Benyahia, Kaicheng Yu, Kamil Bennani-Smires +4

We identify a phenomenon, which we refer to as multi-model forgetting, that occurs when sequentially training multiple deep networks with partially-shared parameters; the performan…

cs.CV2021

Robust Differentiable SVD

Wei Wang, Zheng Dang, Yinlin Hu +2

Eigendecomposition of symmetric matrices is at the heart of many computer vision algorithms. However, the derivatives of the eigenvectors tend to be numerically unstable, whether u…

cs.LG2024

On the Impact of Hard Adversarial Instances on Overfitting in Adversarial Training

Chen Liu, Zhichao Huang, Mathieu Salzmann +2

Adversarial training is a popular method to robustify models against adversarial attacks. However, it exhibits much more severe overfitting than training on clean inputs. In this w…

cs.CV2020

Learning 3D-3D Correspondences for One-shot Partial-to-partial Registration

Zheng Dang, Fei Wang, Mathieu Salzmann

While 3D-3D registration is traditionally tacked by optimization-based methods, recent work has shown that learning-based techniques could achieve faster and more robust results. I…

cs.CV2020

Single-Stage 6D Object Pose Estimation

Yinlin Hu, Pascal Fua, Wei Wang +1

Most recent 6D pose estimation frameworks first rely on a deep network to establish correspondences between 3D object keypoints and 2D image locations and then use a variant of a R…

cs.CV2024

Unlocking Comics: The AI4VA Dataset for Visual Understanding

Peter Grönquist, Deblina Bhattacharjee, Bahar Aydemir +4

In the evolving landscape of deep learning, there is a pressing need for more comprehensive datasets capable of training models across multiple modalities. Concurrently, in digital…

cs.CV2022

Contact-aware Human Motion Forecasting

Wei Mao, Miaomiao Liu, Richard Hartley +1

In this paper, we tackle the task of scene-aware 3D human motion forecasting, which consists of predicting future human poses given a 3D scene and a past human motion. A key challe…

cs.CV2024

3D Single-object Tracking in Point Clouds with High Temporal Variation

Qiao Wu, Kun Sun, Pei An +3

The high temporal variation of the point clouds is the key challenge of 3D single-object tracking (3D SOT). Existing approaches rely on the assumption that the shape variation of t…

cs.CV2023

Linear-Covariance Loss for End-to-End Learning of 6D Pose Estimation

Fulin Liu, Yinlin Hu, Mathieu Salzmann

Most modern image-based 6D object pose estimation methods learn to predict 2D-3D correspondences, from which the pose can be obtained using a PnP solver. Because of the non-differe…

cs.CV2022

Long Term Motion Prediction Using Keyposes

Sena Kiciroglu, Wei Wang, Mathieu Salzmann +1

Long term human motion prediction is essential in safety-critical applications such as human-robot interaction and autonomous driving. In this paper we show that to achieve long te…

cs.CV2026

Coherent and Multi-modality Image Inpainting via Latent Space Optimization

Lingzhi Pan, Tong Zhang, Bingyuan Chen +4

With the advancements in denoising diffusion probabilistic models (DDPMs), image inpainting has significantly evolved from merely filling information based on nearby regions to gen…

cs.CV2020

Robust RGB-based 6-DoF Pose Estimation without Real Pose Annotations

Zhigang Li, Yinlin Hu, Mathieu Salzmann +1

While much progress has been made in 6-DoF object pose estimation from a single RGB image, the current leading approaches heavily rely on real-annotation data. As such, they remain…

cs.CV2025

FastPose-ViT: A Vision Transformer for Real-Time Spacecraft Pose Estimation

Pierre Ancey, Andrew Price, Saqib Javed +1

Estimating the 6-degrees-of-freedom (6DoF) pose of a spacecraft from a single image is critical for autonomous operations like in-orbit servicing and space debris removal. Existing…

cs.CV2021

Attention-based Domain Adaptation for Single Stage Detectors

Vidit Vidit, Mathieu Salzmann

While domain adaptation has been used to improve the performance of object detectors when the training and test data follow different distributions, previous work has mostly focuse…

cs.LG2020

Learning Variations in Human Motion via Mix-and-Match Perturbation

Mohammad Sadegh Aliakbarian, Fatemeh Sadat Saleh, Mathieu Salzmann +3

Human motion prediction is a stochastic process: Given an observed sequence of poses, multiple future motions are plausible. Existing approaches to modeling this stochasticity typi…

cs.CV2018

VIENA2: A Driving Anticipation Dataset

Mohammad Sadegh Aliakbarian, Fatemeh Sadat Saleh, Mathieu Salzmann +3

Action anticipation is critical in scenarios where one needs to react before the action is finalized. This is, for instance, the case in automated driving, where a car needs to, e.…

cs.CV2024

CLOAF: CoLlisiOn-Aware Human Flow

Andrey Davydov, Martin Engilberge, Mathieu Salzmann +1

Even the best current algorithms for estimating body 3D shape and pose yield results that include body self-intersections. In this paper, we present CLOAF, which exploits the diffe…

cs.CV2026

Q-Drift: Quantization-Aware Drift Correction for Diffusion Model Sampling

Sooyoung Ryu, Mathieu Salzmann, Saqib Javed

Post-training quantization (PTQ) is a practical path to deploy large diffusion models, but quantization noise can accumulate over the denoising trajectory and degrade generation qu…

cs.LG2026

Learning to Weight Parameters for Training Data Attribution

Shuangqi Li, Hieu Le, Jingyi Xu +1

We study gradient-based data attribution, aiming to identify which training examples most influence a given output. Existing methods for this task either treat network parameters u…

cs.CV2016

Sample and Filter: Nonparametric Scene Parsing via Efficient Filtering

Mohammad Najafi, Sarah Taghavi Namin, Mathieu Salzmann +1

Scene parsing has attracted a lot of attention in computer vision. While parametric models have proven effective for this task, they cannot easily incorporate new training data. By…

cs.CV2020

GarNet++: Improving Fast and Accurate Static3D Cloth Draping by Curvature Loss

Erhan Gundogdu, Victor Constantin, Shaifali Parashar +4

In this paper, we tackle the problem of static 3D cloth draping on virtual human bodies. We introduce a two-stream deep network model that produces a visually plausible draping of…

cs.CV2018

Statistically Motivated Second Order Pooling

Kaicheng Yu, Mathieu Salzmann

Second-order pooling, a.k.a.~bilinear pooling, has proven effective for deep learning based visual recognition. However, the resulting second-order networks yield a final represent…

cs.CV2016

Shape Interaction Matrix Revisited and Robustified: Efficient Subspace Clustering with Corrupted and Incomplete Data

Pan Ji, Mathieu Salzmann, Hongdong Li

The Shape Interaction Matrix (SIM) is one of the earliest approaches to performing subspace clustering (i.e., separating points drawn from a union of subspaces). In this paper, we…

cs.CV2015

Kernel Methods on Riemannian Manifolds with Gaussian RBF Kernels

Sadeep Jayasumana, Richard Hartley, Mathieu Salzmann +2

In this paper, we develop an approach to exploiting kernel methods with manifold-valued data. In many computer vision problems, the data can be naturally represented as points on a…

cs.LG2025

QT-DoG: Quantization-aware Training for Domain Generalization

Saqib Javed, Hieu Le, Mathieu Salzmann

A key challenge in Domain Generalization (DG) is preventing overfitting to source domains, which can be mitigated by finding flatter minima in the loss landscape. In this work, we…

cs.CV2020

Temporally-Transferable Perturbations: Efficient, One-Shot Adversarial Attacks for Online Visual Object Trackers

Krishna Kanth Nakka, Mathieu Salzmann

In recent years, the trackers based on Siamese networks have emerged as highly effective and efficient for visual object tracking (VOT). While these methods were shown to be vulner…

cs.CV2024

Generalize or Detect? Towards Robust Semantic Segmentation Under Multiple Distribution Shifts

Zhitong Gao, Bingnan Li, Mathieu Salzmann +1

In open-world scenarios, where both novel classes and domains may exist, an ideal segmentation model should detect anomaly classes for safety and generalize to new domains. However…

cs.CV2024

SINDER: Repairing the Singular Defects of DINOv2

Haoqi Wang, Tong Zhang, Mathieu Salzmann

Vision Transformer models trained on large-scale datasets, although effective, often exhibit artifacts in the patch token they extract. While such defects can be alleviated by re-t…

cs.CV2023

3D-Aware Hypothesis & Verification for Generalizable Relative Object Pose Estimation

Chen Zhao, Tong Zhang, Mathieu Salzmann

Prior methods that tackle the problem of generalizable object pose estimation highly rely on having dense views of the unseen object. By contrast, we address the scenario where onl…

cs.LG2024

Controlling the Fidelity and Diversity of Deep Generative Models via Pseudo Density

Shuangqi Li, Chen Liu, Tong Zhang +3

We introduce an approach to bias deep generative models, such as GANs and diffusion models, towards generating data with either enhanced fidelity or increased diversity. Our approa…

cs.CV2022

Leverage Your Local and Global Representations: A New Self-Supervised Learning Strategy

Tong Zhang, Congpei Qiu, Wei Ke +2

Self-supervised learning (SSL) methods aim to learn view-invariant representations by maximizing the similarity between the features extracted from different crops of the same imag…

cs.CV2018

Beyond One Glance: Gated Recurrent Architecture for Hand Segmentation

Wei Wang, Kaicheng Yu, Joachim Hugonot +2

As mixed reality is gaining increased momentum, the development of effective and efficient solutions to egocentric hand segmentation is becoming critical. Traditional segmentation…

cs.CV2017

Soft Correspondences in Multimodal Scene Parsing

Sarah Taghavi Namin, Mohammad Najafi, Mathieu Salzmann +1

Exploiting multiple modalities for semantic scene parsing has been shown to improve accuracy over the singlemodality scenario. However multimodal datasets often suffer from problem…

cs.LG2020

On the Loss Landscape of Adversarial Training: Identifying Challenges and How to Overcome Them

Chen Liu, Mathieu Salzmann, Tao Lin +2

We analyze the influence of adversarial training on the loss landscape of machine learning models. To this end, we first provide analytical studies of the properties of adversarial…

cs.CV2016

Semantic-Aware Depth Super-Resolution in Outdoor Scenes

Miaomiao Liu, Mathieu Salzmann, Xuming He

While depth sensors are becoming increasingly popular, their spatial resolution often remains limited. Depth super-resolution therefore emerged as a solution to this problem. Despi…

cs.CV2015

When VLAD met Hilbert

Mehrtash Harandi, Mathieu Salzmann, Fatih Porikli

Vectors of Locally Aggregated Descriptors (VLAD) have emerged as powerful image/video representations that compete with or even outperform state-of-the-art approaches on many chall…

cs.CV2022

Dyadic Human Motion Prediction

Isinsu Katircioglu, Costa Georgantas, Mathieu Salzmann +1

Prior work on human motion forecasting has mostly focused on predicting the future motion of single subjects in isolation from their past pose sequence. In the presence of closely…

cs.CV2024

TempSAL -- Uncovering Temporal Information for Deep Saliency Prediction

Bahar Aydemir, Ludo Hoffstetter, Tong Zhang +2

Deep saliency prediction algorithms complement the object recognition features, they typically rely on additional information, such as scene context, semantic relationships, gaze d…

cs.LG2025

Towards Self-Supervised Covariance Estimation in Deep Heteroscedastic Regression

Megh Shukla, Aziz Shameem, Mathieu Salzmann +1

Deep heteroscedastic regression models the mean and covariance of the target distribution through neural networks. The challenge arises from heteroscedasticity, which implies that…

cs.LG2020

Contextually Plausible and Diverse 3D Human Motion Prediction

Sadegh Aliakbarian, Fatemeh Sadat Saleh, Lars Petersson +2

We tackle the task of diverse 3D human motion prediction, that is, forecasting multiple plausible future 3D poses given a sequence of observed 3D poses. In this context, a popular…

cs.CV2014

Iteratively Reweighted Graph Cut for Multi-label MRFs with Non-convex Priors

Thalaiyasingam Ajanthan, Richard Hartley, Mathieu Salzmann +1

While widely acknowledged as highly effective in computer vision, multi-label MRFs with non-convex priors are difficult to optimize. To tackle this, we introduce an algorithm that…

cs.LG2026

LoRaQ: Optimized Low Rank Approximation for 4-bit Quantization

Yann Bouquet, Alireza Khodamoradi, Sophie Yáng Shen +2

Post-training quantization (PTQ) is essential for deploying large diffusion transformers on resource-constrained hardware, but aggressive 4-bit quantization significantly degrades…

cs.CV2019

Neural Scene Decomposition for Multi-Person Motion Capture

Helge Rhodin, Victor Constantin, Isinsu Katircioglu +2

Learning general image representations has proven key to the success of many computer vision tasks. For example, many approaches to image understanding problems rely on deep networ…

cs.CV2019

Recurrent U-Net for Resource-Constrained Segmentation

Wei Wang, Kaicheng Yu, Joachim Hugonot +2

State-of-the-art segmentation methods rely on very deep networks that are not always easy to train without very large training datasets and tend to be relatively slow to run on sta…

astro-ph.IM2025

Deep learning to improve the discovery of near-Earth asteroids in the Zwicky Transient Facility

Belén Yu Irureta-Goyena, George Helou, Jean-Paul Kneib +7

We present a novel pipeline that uses a convolutional neural network (CNN) to improve the detection capability of near-Earth asteroids (NEAs) in the context of planetary defense. O…