papers

Publications (48)

cs.CV2016

Do We Really Need to Collect Millions of Faces for Effective Face Recognition?

Iacopo Masi, Anh Tuan Tran, Jatuporn Toy Leksut +2

Face recognition capabilities have recently made extraordinary leaps. Though this progress is at least partially due to ballooning training set sizes -- huge numbers of face images…

cs.CV2023

A Whac-A-Mole Dilemma: Shortcuts Come in Multiples Where Mitigating One Amplifies Others

Zhiheng Li, Ivan Evtimov, Albert Gordo +5

Machine learning models have been found to learn shortcuts -- unintended decision rules that are unable to generalize -- undermining models' reliability. Previous works address thi…

cs.CV2013

Single View Depth Estimation from Examples

Tal Hassner, Ronen Basri

We describe a non-parametric, "example-based" method for estimating the depth of an object, viewed in a single photo. Our method consults a database of example 3D geometries, searc…

cs.CV2019

Precise Detection in Densely Packed Scenes

Eran Goldman, Roei Herzig, Aviv Eisenschtat +4

Man-made scenes can be densely packed, containing numerous objects, often identical, positioned in close proximity. We show that precise object detection in such scenes remains a c…

cs.CV2019

Learn Stereo, Infer Mono: Siamese Networks for Self-Supervised, Monocular, Depth Estimation

Matan Goldman, Tal Hassner, Shai Avidan

The field of self-supervised monocular depth estimation has seen huge advancements in recent years. Most methods assume stereo data is available during training but usually under-u…

cs.CV2022

Task Grouping for Multilingual Text Recognition

Jing Huang, Kevin J Liang, Rama Kovvuri +1

Most existing OCR methods focus on alphanumeric characters due to the popularity of English and numbers, as well as their corresponding datasets. On extending the characters to mor…

cs.CV2021

img2pose: Face Alignment and Detection via 6DoF, Face Pose Estimation

Vítor Albiero, Xingyu Chen, Xi Yin +2

We propose real-time, six degrees of freedom (6DoF), 3D face pose estimation without face detection or landmark localization. We observe that estimating the 6DoF rigid transformati…

cs.CV2021

HyperSeg: Patch-wise Hypernetwork for Real-time Semantic Segmentation

Yuval Nirkin, Lior Wolf, Tal Hassner

We present a novel, real-time, semantic segmentation network in which the encoder both encodes and generates the parameters (weights) of the decoder. Furthermore, to allow maximal…

cs.CV2017

On Face Segmentation, Face Swapping, and Face Perception

Yuval Nirkin, Iacopo Masi, Anh Tuan Tran +2

We show that even when face images are unconstrained and arbitrarily paired, face swapping between them is actually quite simple. To this end, we make the following contributions.…

cs.CV2014

Dense Correspondences Across Scenes and Scales

Moria Tau, Tal Hassner

We seek a practical method for establishing dense correspondences between two images with similar content, but possibly different 3D scenes. One of the challenges in designing such…

cs.LG2023

HyperMix: Out-of-Distribution Detection and Classification in Few-Shot Settings

Nikhil Mehta, Kevin J Liang, Jing Huang +3

Out-of-distribution (OOD) detection is an important topic for real-world machine learning systems, but settings with limited in-distribution samples have been underexplored. Such f…

cs.LG2021

Single Layer Predictive Normalized Maximum Likelihood for Out-of-Distribution Detection

Koby Bibas, Meir Feder, Tal Hassner

Detecting out-of-distribution (OOD) samples is vital for developing machine learning based models for critical safety systems. Common approaches for OOD detection assume access to…

cs.CV2015

GPU-Based Computation of 2D Least Median of Squares with Applications to Fast and Robust Line Detection

Gil Shapira, Tal Hassner

The 2D Least Median of Squares (LMS) is a popular tool in robust regression because of its high breakdown point: up to half of the input data can be contaminated with outliers with…

cs.CV2023

Reverse Engineering of Generative Models: Inferring Model Hyperparameters from Generated Images

Vishal Asnani, Xi Yin, Tal Hassner +1

State-of-the-art (SOTA) Generative Models (GMs) can synthesize photo-realistic images that are hard for humans to distinguish from genuine photos. Identifying and understanding man…

cs.CV2020

DeepFake Detection Based on the Discrepancy Between the Face and its Context

Yuval Nirkin, Lior Wolf, Yosi Keller +1

We propose a method for detecting face swapping and other identity manipulations in single images. Face swapping methods, such as DeepFake, manipulate the face region, aiming to ad…

cs.CL2024

Navigating Text-to-Image Generative Bias across Indic Languages

Surbhi Mittal, Arnav Sudan, Mayank Vatsa +3

This research investigates biases in text-to-image (TTI) models for the Indic languages widely spoken across India. It evaluates and compares the generative performance and cultura…

cs.CV2021

TextStyleBrush: Transfer of Text Aesthetics from a Single Example

Praveen Krishnan, Rama Kovvuri, Guan Pang +2

We present a novel approach for disentangling the content of a text image from all aspects of its appearance. The appearance representation we derive can then be applied to new con…

cs.CV2016

Face Recognition Using Deep Multi-Pose Representations

Wael AbdAlmageed, Yue Wua, Stephen Rawlsa +9

We introduce our method and system for face recognition using multiple pose-aware deep learning models. In our representation, a face image is processed by several pose-specific de…

cs.LG2023

Simple Transferability Estimation for Regression Tasks

Cuong N. Nguyen, Phong Tran, Lam Si Tung Ho +4

We consider transferability estimation, the problem of estimating how well deep learning models transfer from a source to a target task. We focus on regression tasks, which receive…

cs.CV2022

Few-shot Learning with Noisy Labels

Kevin J Liang, Samrudhdhi B. Rangrej, Vladan Petrovic +1

Few-shot learning (FSL) methods typically assume clean support sets with accurately labeled samples when training on novel classes. This assumption can often be unrealistic: suppor…

cs.CV2016

Regressing Robust and Discriminative 3D Morphable Models with a very Deep Neural Network

Anh Tuan Tran, Tal Hassner, Iacopo Masi +1

The 3D shapes of faces are well known to be discriminative. Yet despite this, they are rarely used for face recognition and always under controlled viewing conditions. We claim tha…

cs.CV2022

FSGANv2: Improved Subject Agnostic Face Swapping and Reenactment

Yuval Nirkin, Yosi Keller, Tal Hassner

We present Face Swapping GAN (FSGAN) for face swapping and reenactment. Unlike previous work, we offer a subject agnostic swapping scheme that can be applied to pairs of faces with…

cs.CV2016

Facial Landmark Detection with Tweaked Convolutional Neural Networks

Yue Wu, Tal Hassner, KangGeon Kim +2

We present a novel convolutional neural network (CNN) design for facial landmark coordinate regression. We examine the intermediate features of a standard CNN trained for landmark…

cs.CV2016

Pooling Faces: Template based Face Recognition with Pooled Face Images

Tal Hassner, Iacopo Masi, Jungyeon Kim +4

We propose a novel approach to template based face recognition. Our dual goal is to both increase recognition accuracy and reduce the computational and storage costs of template ma…

cs.CV2020

Mask TextSpotter v3: Segmentation Proposal Network for Robust Scene Text Spotting

Minghui Liao, Guan Pang, Jing Huang +2

Recent end-to-end trainable methods for scene text spotting, integrating detection and recognition, showed much progress. However, most of the current arbitrary-shape scene text sp…

cs.CV2025

Fine-Grained Erasure in Text-to-Image Diffusion-based Foundation Models

Kartik Thakral, Tamar Glaser, Tal Hassner +2

Existing unlearning algorithms in text-to-image generative models often fail to preserve the knowledge of semantically related concepts when removing specific target concepts: a ch…

cs.CV2016

The CUDA LATCH Binary Descriptor: Because Sometimes Faster Means Better

Christopher Parker, Matthew Daiter, Kareem Omar +2

Accuracy, descriptor size, and the time required for extraction and matching are all important factors when selecting local image descriptors. To optimize over all these requiremen…

cs.LG2022

Generalization Bounds for Deep Transfer Learning Using Majority Predictor Accuracy

Cuong N. Nguyen, Lam Si Tung Ho, Vu Dinh +2

We analyze new generalization bounds for deep learning models trained by transfer learning from a source to a target task. Our bounds utilize a quantity called the majority predict…

cs.LG2019

Transferability and Hardness of Supervised Classification Tasks

Anh T. Tran, Cuong V. Nguyen, Tal Hassner

We propose a novel approach for estimating the difficulty and transferability of supervised classification tasks. Unlike previous work, our approach is solution agnostic and does n…

cs.CV2015

Wide baseline stereo matching with convex bounded-distortion constraints

Meirav Galun, Tal Amir, Tal Hassner +2

Finding correspondences in wide baseline setups is a challenging problem. Existing approaches have focused largely on developing better feature descriptors for correspondence and o…

cs.CV2017

FacePoseNet: Making a Case for Landmark-Free Face Alignment

Fengju Chang, Anh Tuan Tran, Tal Hassner +3

We show how a simple convolutional neural network (CNN) can be trained to accurately and robustly regress 6 degrees of freedom (6DoF) 3D head pose, directly from image intensities.…

cs.CV2023

MaLP: Manipulation Localization Using a Proactive Scheme

Vishal Asnani, Xi Yin, Tal Hassner +1

Advancements in the generation quality of various Generative Models (GMs) has made it necessary to not only perform binary manipulation detection but also localize the modified pix…

cs.LG2020

LEEP: A New Measure to Evaluate Transferability of Learned Representations

Cuong V. Nguyen, Tal Hassner, Matthias Seeger +1

We introduce a new measure to evaluate the transferability of representations learned by classifiers. Our measure, the Log Expected Empirical Prediction (LEEP), is simple and easy…

cs.CV2018

ExpNet: Landmark-Free, Deep, 3D Facial Expressions

Feng-Ju Chang, Anh Tuan Tran, Tal Hassner +3

We describe a deep learning based method for estimating 3D facial expression coefficients. Unlike previous work, our process does not relay on facial landmark detection methods as…

cs.CV2017

Temporal Tessellation: A Unified Approach for Video Analysis

Dotan Kaufman, Gil Levi, Tal Hassner +1

We present a general approach to video understanding, inspired by semantic transfer techniques that have been successfully used for 2D image analysis. Our method considers a video…

cs.CV2025

Continual Unlearning for Foundational Text-to-Image Models without Generalization Erosion

Kartik Thakral, Tamar Glaser, Tal Hassner +2

How can we effectively unlearn selected concepts from pre-trained generative foundation models without resorting to extensive retraining? This research introduces `continual unlear…

cs.CV2022

Proactive Image Manipulation Detection

Vishal Asnani, Xi Yin, Tal Hassner +2

Image manipulation detection algorithms are often trained to discriminate between images manipulated with particular Generative Models (GMs) and genuine/real images, yet generalize…

cs.CV2023

GliTr: Glimpse Transformers with Spatiotemporal Consistency for Online Action Prediction

Samrudhdhi B Rangrej, Kevin J Liang, Tal Hassner +1

Many online action prediction models observe complete frames to locate and attend to informative subregions in the frames called glimpses and recognize an ongoing action based on g…

cs.CV2018

Extreme 3D Face Reconstruction: Seeing Through Occlusions

Anh Tuan Tran, Tal Hassner, Iacopo Masi +3

Existing single view, 3D face reconstruction methods can produce beautifully detailed 3D results, but typically only for near frontal, unobstructed viewpoints. We describe a system…

cs.CV2022

You Only Need a Good Embeddings Extractor to Fix Spurious Correlations

Raghav Mehta, Vítor Albiero, Li Chen +4

Spurious correlations in training data often lead to robustness issues since models learn to use them as shortcuts. For example, when predicting whether an object is a cow, a model…

cs.CV2019

FSGAN: Subject Agnostic Face Swapping and Reenactment

Yuval Nirkin, Yosi Keller, Tal Hassner

We present Face Swapping GAN (FSGAN) for face swapping and reenactment. Unlike previous work, FSGAN is subject agnostic and can be applied to pairs of faces without requiring train…

cs.LG2019

Toward Understanding Catastrophic Forgetting in Continual Learning

Cuong V. Nguyen, Alessandro Achille, Michael Lam +3

We study the relationship between catastrophic forgetting and properties of task sequences. In particular, given a sequence of tasks, we would like to understand which properties o…

cs.LG2024

On Responsible Machine Learning Datasets with Fairness, Privacy, and Regulatory Norms

Surbhi Mittal, Kartik Thakral, Richa Singh +4

Artificial Intelligence (AI) has made its way into various scientific fields, providing astonishing improvements over existing algorithms for a wide variety of tasks. In recent yea…

cs.CV2014

Effective Face Frontalization in Unconstrained Images

Tal Hassner, Shai Harel, Eran Paz +1

"Frontalization" is the process of synthesizing frontal facing views of faces appearing in single unconstrained photos. Recent reports have suggested that this process may substant…

cs.CV2021

A Multiplexed Network for End-to-End, Multilingual OCR

Jing Huang, Guan Pang, Rama Kovvuri +5

Recent advances in OCR have shown that an end-to-end (E2E) training pipeline that includes both detection and recognition leads to the best results. However, many existing methods…

cs.CV2019

Balancing Specialization, Generalization, and Compression for Detection and Tracking

Dotan Kaufman, Koby Bibas, Eran Borenstein +2

We propose a method for specializing deep detectors and trackers to restricted settings. Our approach is designed with the following goals in mind: (a) Improving accuracy in restri…

cs.CV2015

LATCH: Learned Arrangements of Three Patch Codes

Gil Levi, Tal Hassner

We present a novel means of describing local image appearances using binary strings. Binary descriptors have drawn increasing interest in recent years due to their speed and low me…

cs.CV2021

TextOCR: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text

Amanpreet Singh, Guan Pang, Mandy Toh +3

A crucial component for the scene text based reasoning required for TextVQA and TextCaps datasets involve detecting and recognizing text present in the images using an optical char…