FaceNet: A Unified Embedding for Face Recognition and Clustering
arXiv:1503.03832 · doi:10.1109/CVPR.2015.7298682
Abstract
Despite significant recent advances in the field of face recognition, implementing face verification and recognition efficiently at scale presents serious challenges to current approaches. In this paper we present a system, called FaceNet, that directly learns a mapping from face images to a compact Euclidean space where distances directly correspond to a measure of face similarity. Once this space has been produced, tasks such as face recognition, verification and clustering can be easily implemented using standard techniques with FaceNet embeddings as feature vectors. Our method uses a deep convolutional network trained to directly optimize the embedding itself, rather than an intermediate bottleneck layer as in previous deep learning approaches. To train, we use triplets of roughly aligned matching / non-matching face patches generated using a novel online triplet mining method. The benefit of our approach is much greater representational efficiency: we achieve state-of-the-art face recognition performance using only 128-bytes per face. On the widely used Labeled Faces in the Wild (LFW) dataset, our system achieves a new record accuracy of 99.63%. On YouTube Faces DB it achieves 95.12%. Our system cuts the error rate in comparison to the best published result by 30% on both datasets. We also introduce the concept of harmonic embeddings, and a harmonic triplet loss, which describe different versions of face embeddings (produced by different networks) that are compatible to each other and allow for direct comparison between each other.
Also published, in Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition 2015
Cited by in corpus (458)
- Deep Facial Expression Recognition: A Survey
- Additive Margin Softmax for Face Verification
- Graph Contrastive Learning with Adaptive Augmentation
- Deep Face Recognition: A Survey
- Contrastive Representation Learning: A Framework and Review
- A Strong Baseline and Batch Normalization Neck for Deep Person Re-identification
- Deep Isolation Forest for Anomaly Detection
- End-to-End Comparative Attention Networks for Person Re-identification
- Deep Learning for Deepfakes Creation and Detection: A Survey
- Trunk-Branch Ensemble Convolutional Neural Networks for Video-based Face Recognition
- Image-based 3D Object Reconstruction: State-of-the-Art and Trends in the Deep Learning Era
- 3DFeat-Net: Weakly Supervised Local 3D Features for Point Cloud Registration
- Facial Landmark Detection: a Literature Survey
- SurfaceNet: An End-to-end 3D Neural Network for Multiview Stereopsis
- Deep Cosine Metric Learning for Person Re-Identification
- GANFIT: Generative Adversarial Network Fitting for High Fidelity 3D Face Reconstruction
- Each Part Matters: Local Patterns Facilitate Cross-view Geo-localization
- AdvHat: Real-world adversarial attack on ArcFace Face ID system
- Loss Functions and Metrics in Deep Learning
- Multi-Task Convolutional Neural Network for Pose-Invariant Face Recognition
- Parameter Sharing Exploration and Hetero-Center based Triplet Loss for Visible-Thermal Person Re-Identification
- Deep Attributes Driven Multi-Camera Person Re-identification
- On the Reconstruction of Face Images from Deep Face Templates
- CIAGAN: Conditional Identity Anonymization Generative Adversarial Networks
- Low-resolution Face Recognition in the Wild via Selective Knowledge Distillation
- Strengths and Weaknesses of Deep Learning Models for Face Recognition Against Image Degradations
- SkeleMotion: A New Representation of Skeleton Joint Sequences Based on Motion Information for 3D Action Recognition
- Demographic Bias in Biometrics: A Survey on an Emerging Challenge
- SegMap: Segment-based mapping and localization using data-driven descriptors
- Exploiting Unlabeled Data in CNNs by Self-supervised Learning to Rank
- MIN2Net: End-to-End Multi-Task Learning for Subject-Independent Motor Imagery EEG Classification
- On Low-Resolution Face Recognition in the Wild: Comparisons and New Techniques
- Towards Transferable Adversarial Attack against Deep Face Recognition
- Representation Learning by Rotating Your Faces
- GraphIQA: Learning Distortion Graph Representations for Blind Image Quality Assessment
- Triplet Probabilistic Embedding for Face Verification and Clustering
- Discrimination-aware Network Pruning for Deep Model Compression
- Self-Supervised Learning for Videos: A Survey
- Palmprint Recognition in Uncontrolled and Uncooperative Environment
- Exploring Spatial Significance via Hybrid Pyramidal Graph Network for Vehicle Re-identification
- SFace: Sigmoid-Constrained Hypersphere Loss for Robust Face Recognition
- A Comprehensive Performance Evaluation of Deformable Face Tracking "In-the-Wild"
- An End to End Deep Neural Network for Iris Segmentation in Unconstraint Scenarios
- Recent Advances in Deep Learning Techniques for Face Recognition
- A Survey on Face Data Augmentation
- Relational Reflection Entity Alignment
- 3D Convolutional Neural Networks for Cross Audio-Visual Matching Recognition
- Self Paced Deep Learning for Weakly Supervised Object Detection
- An IoT Endpoint System-on-Chip for Secure and Energy-Efficient Near-Sensor Analytics
- Time Series Change Point Detection with Self-Supervised Contrastive Predictive Coding
- Joint Face Alignment and 3D Face Reconstruction with Application to Face Recognition
- Boosting the Speed of Entity Alignment 10*: Dual Attention Matching Network with Normalized Hard Sample Mining
- Personalized Age Progression with Bi-level Aging Dictionary Learning
- Multi-modal Multi-channel Target Speech Separation
- CMTR: Cross-modality Transformer for Visible-infrared Person Re-identification
- Contrast-Phys: Unsupervised Video-based Remote Physiological Measurement via Spatiotemporal Contrast
- Deep learning and face recognition: the state of the art
- Self-supervised remote sensing feature learning: Learning Paradigms, Challenges, and Future Works
- Transferring Knowledge Fragments for Learning Distance Metric from A Heterogeneous Domain
- Visual Search at Alibaba
- Few-Shot Deep Adversarial Learning for Video-based Person Re-identification
- Self-Supervised Contrastive Learning for Code Retrieval and Summarization via Semantic-Preserving Transformations
- Knowledge-Preserving Incremental Social Event Detection via Heterogeneous GNNs
- Enhancing Remote Sensing Image Retrieval with Triplet Deep Metric Learning Network
- Hierarchy-based Image Embeddings for Semantic Image Retrieval
- How Image Degradations Affect Deep CNN-based Face Recognition?
- Deep Generalized Max Pooling
- Incomplete Descriptor Mining with Elastic Loss for Person Re-Identification
- Physical Adversarial Attack meets Computer Vision: A Decade Survey
- Face Clustering: Representation and Pairwise Constraints
- Deep Hashing for Secure Multimodal Biometrics
- Metric-Learning based Deep Hashing Network for Content Based Retrieval of Remote Sensing Images
- Families in the Wild (FIW): Large-Scale Kinship Image Database and Benchmarks
- Part-based Deep Hashing for Large-scale Person Re-identification
- Margin Preserving Self-paced Contrastive Learning Towards Domain Adaptation for Medical Image Segmentation
- Automated Radiology Report Generation: A Review of Recent Advances
- Deep Learning Based Single Sample Per Person Face Recognition: A Survey
- Deep Graph Matching via Blackbox Differentiation of Combinatorial Solvers
- Deep Convolutional Neural Networks in the Face of Caricature: Identity and Image Revealed
- Learning View-Specific Deep Networks for Person Re-Identification
- Lightweight Multi-Branch Network for Person Re-Identification
- Learning Multi-Attention Context Graph for Group-Based Re-Identification
- Learning from Millions of 3D Scans for Large-scale 3D Face Recognition
- A Practical Contrastive Learning Framework for Single-Image Super-Resolution
- SleePyCo: Automatic Sleep Scoring with Feature Pyramid and Contrastive Learning
- Significance of Softmax-based Features in Comparison to Distance Metric Learning-based Features
- Deep learning on butterfly phenotypes tests evolution's oldest mathematical model
- Meta Balanced Network for Fair Face Recognition
- A Similarity Measure for Material Appearance
- Optimizing Rank-based Metrics with Blackbox Differentiation
- Stacking-Based Deep Neural Network: Deep Analytic Network for Pattern Classification
- Towards End-to-End Face Recognition through Alignment Learning
- Contrast-Phys+: Unsupervised and Weakly-supervised Video-based Remote Physiological Measurement via Spatiotemporal Contrast
- TripletTrack: 3D Object Tracking using Triplet Embeddings and LSTM
- Fast-GANFIT: Generative Adversarial Network for High Fidelity 3D Face Reconstruction
- REVAMPT: Real-time Edge Video Analytics for Multi-camera Privacy-aware Pedestrian Tracking
- Biometric Template Protection for Neural-Network-based Face Recognition Systems: A Survey of Methods and Evaluation Techniques
- A Review of Predictive and Contrastive Self-supervised Learning for Medical Images
- Average Biased ReLU Based CNN Descriptor for Improved Face Retrieval
- Multi-Fold Gabor, PCA and ICA Filter Convolution Descriptor for Face Recognition
- Computer Vision on X-ray Data in Industrial Production and Security Applications: A Comprehensive Survey
- Human-Centric Multimodal Machine Learning: Recent Advances and Testbed on AI-based Recruitment
- Semi-Cycled Generative Adversarial Networks for Real-World Face Super-Resolution
- Are GAN-based Morphs Threatening Face Recognition?
- Line as a Visual Sentence: Context-aware Line Descriptor for Visual Localization
- Group Re-Identification with Multi-grained Matching and Integration
- Feature Completion Transformer for Occluded Person Re-identification
- Geo-Localization via Ground-to-Satellite Cross-View Image Retrieval
- APANet: Adaptive Prototypes Alignment Network for Few-Shot Semantic Segmentation
- Towards NIR-VIS Masked Face Recognition
- Learning Efficient Representations for Keyword Spotting with Triplet Loss
- Multiscale CNN based Deep Metric Learning for Bioacoustic Classification: Overcoming Training Data Scarcity Using Dynamic Triplet Loss
- StyleCariGAN: Caricature Generation via StyleGAN Feature Map Modulation
- SAAN: Similarity-aware attention flow network for change detection with VHR remote sensing images
- Relational Deep Feature Learning for Heterogeneous Face Recognition
- EmoBed: Strengthening Monomodal Emotion Recognition via Training with Crossmodal Emotion Embeddings
- Image-based model parameter optimization using Model-Assisted Generative Adversarial Networks
- Local Directional Relation Pattern for Unconstrained and Robust Face Retrieval
- GCoNet+: A Stronger Group Collaborative Co-Salient Object Detector
- Self-supervised Audiovisual Representation Learning for Remote Sensing Data
- PoreNet: CNN-based Pore Descriptor for High-resolution Fingerprint Recognition
- Fine-Grained Fashion Similarity Prediction by Attribute-Specific Embedding Learning
- Deep Convolutional Neural Network Features and the Original Image
- Semi-supervised Adversarial Learning to Generate Photorealistic Face Images of New Identities from 3D Morphable Model
- Towards Unconstrained Palmprint Recognition on Consumer Devices: a Literature Review
- Deep Learning-Based Robust Multi-Object Tracking via Fusion of mmWave Radar and Camera Sensors
- No-Reference Image Quality Assessment by Hallucinating Pristine Features
- Sampling Strategies for GAN Synthetic Data
- Neural Multi-Atlas Label Fusion: Application to Cardiac MR Images
- Anchor Cascade for Efficient Face Detection
- Biometric Backdoors: A Poisoning Attack Against Unsupervised Template Updating
- Augmenting Visual Place Recognition with Structural Cues
- Multi-Scale Thermal to Visible Face Verification via Attribute Guided Synthesis
- OPOM: Customized Invisible Cloak towards Face Privacy Protection
- Improving Deep Facial Phenotyping for Ultra-rare Disorder Verification Using Model Ensembles
- Polarimetric Thermal to Visible Face Verification via Attribute Preserved Synthesis
- Are Out-of-Distribution Detection Methods Effective on Large-Scale Datasets?
- Universal Multimodal Representation for Language Understanding
- Fisher Discriminant Triplet and Contrastive Losses for Training Siamese Networks
- Unveiling Energy Efficiency in Deep Learning: Measurement, Prediction, and Scoring across Edge Devices
- Learning Contrastive Self-Distillation for Ultra-Fine-Grained Visual Categorization Targeting Limited Samples
- Adversarial Privacy-preserving Filter
- Adaptive neighborhood Metric learning
- Context-self contrastive pretraining for crop type semantic segmentation
- Batch Coherence-Driven Network for Part-aware Person Re-Identification
- Unsupervised Learning of Deep Features for Music Segmentation
- Not Only Generative Art: Stable Diffusion for Content-Style Disentanglement in Art Analysis
- Disentangling Semantic-to-visual Confusion for Zero-shot Learning
- DACAD: Domain Adaptation Contrastive Learning for Anomaly Detection in Multivariate Time Series
- Entropic Out-of-Distribution Detection: Seamless Detection of Unknown Examples
- Simultaneous multi-view instance detection with learned geometric soft-constraints
- Multi-view Incremental Segmentation of 3D Point Clouds for Mobile Robots
- M-VAD Names: a Dataset for Video Captioning with Naming
- Artificial intelligence for topic modelling in Hindu philosophy: mapping themes between the Upanishads and the Bhagavad Gita
- Vehicle Attribute Recognition by Appearance: Computer Vision Methods for Vehicle Type, Make and Model Classification
- Low-Resolution Object Recognition with Cross-Resolution Relational Contrastive Distillation
- Hierarchical Skeleton Meta-Prototype Contrastive Learning with Hard Skeleton Mining for Unsupervised Person Re-Identification
- CoLES: Contrastive Learning for Event Sequences with Self-Supervision
- Constrained Design of Deep Iris Networks
- Survey on the Analysis and Modeling of Visual Kinship: A Decade in the Making
- Self-supervised Multi-view Person Association and Its Applications
- Graph-based Visual-Semantic Entanglement Network for Zero-shot Image Recognition
- Sampling Agnostic Feature Representation for Long-Term Person Re-identification
- Exploration of Class Center for Fine-Grained Visual Classification
- Multi-Branch Deep Radial Basis Function Networks for Facial Emotion Recognition
- A Feature-space Multimodal Data Augmentation Technique for Text-video Retrieval
- Heri-Graphs: A Workflow of Creating Datasets for Multi-modal Machine Learning on Graphs of Heritage Values and Attributes with Social Media
- Locality-aware Channel-wise Dropout for Occluded Face Recognition
- Graph Jigsaw Learning for Cartoon Face Recognition
- Achieving Better Kinship Recognition Through Better Baseline
- Regular Polytope Networks
- MixNet: Joining Force of Classical and Modern Approaches Toward the Comprehensive Pipeline in Motor Imagery EEG Classification
- Large Scale Landmark Recognition via Deep Metric Learning
- Enhancing Interactive Image Retrieval With Query Rewriting Using Large Language Models and Vision Language Models
- Style Alignment based Dynamic Observation Method for UAV-View Geo-localization
- Lifelong Adaptive Machine Learning for Sensor-based Human Activity Recognition Using Prototypical Networks
- Sharing Matters for Generalization in Deep Metric Learning
- Multi-channel learning for integrating structural hierarchies into context-dependent molecular representation
- SFA: Small Faces Attention Face Detector
- Unsupervised Deep Metric Learning via Orthogonality based Probabilistic Loss
- Physical Adversarial Attacks for Surveillance: A Survey
- OA-Mine: Open-World Attribute Mining for E-Commerce Products with Weak Supervision
- Offline versus Online Triplet Mining based on Extreme Distances of Histopathology Patches
- LOTR: Face Landmark Localization Using Localization Transformer
- Improving Collaborative Metric Learning with Efficient Negative Sampling
- Video-to-Music Recommendation using Temporal Alignment of Segments
- Simultaneous Subspace Clustering and Cluster Number Estimating based on Triplet Relationship
- Contrastive Cross-Modal Pre-Training: A General Strategy for Small Sample Medical Imaging
- Dolphin: A Spoken Language Proficiency Assessment System for Elementary Education
- Quantum Transfer Learning with Adversarial Robustness for Classification of High-Resolution Image Datasets
- Deep Heterogeneous Hashing for Face Video Retrieval
- AirLoop: Lifelong Loop Closure Detection
- FKIMNet: A Finger Dorsal Image Matching Network Comparing Component (Major, Minor and Nail) Matching with Holistic (Finger Dorsal) Matching
- MilliTRACE-IR: Contact Tracing and Temperature Screening via mm-Wave and Infrared Sensing
- Identity-preserving Face Recovery from Stylized Portraits
- Emotion Separation and Recognition from a Facial Expression by Generating the Poker Face with Vision Transformers
- FlakyFix: Using Large Language Models for Predicting Flaky Test Fix Categories and Test Code Repair
- View-Invariant, Occlusion-Robust Probabilistic Embedding for Human Pose
- Leveraging Diffusion For Strong and High Quality Face Morphing Attacks
- NuPS: A Parameter Server for Machine Learning with Non-Uniform Parameter Access
- Unsupervised learning from videos using temporal coherency deep networks
- BC-GAN: A Generative Adversarial Network for Synthesizing a Batch of Collocated Clothing
- Black-Box Face Recovery from Identity Features
- Controllable Generation with Text-to-Image Diffusion Models: A Survey
- Deep Efficient Continuous Manifold Learning for Time Series Modeling
- Optimizing speed/accuracy trade-off for person re-identification via knowledge distillation
- Age-Oriented Face Synthesis with Conditional Discriminator Pool and Adversarial Triplet Loss
- Regional Attention Network (RAN) for Head Pose and Fine-grained Gesture Recognition
- Feature Re-Learning with Data Augmentation for Video Relevance Prediction
- Sequential convolutional network for behavioral pattern extraction in gait recognition
- Rank-Consistency Deep Hashing for Scalable Multi-Label Image Search
- Facial UV Map Completion for Pose-invariant Face Recognition: A Novel Adversarial Approach based on Coupled Attention Residual UNets
- Crossing You in Style: Cross-modal Style Transfer from Music to Visual Arts
- Avoiding Echo-Responses in a Retrieval-Based Conversation System
- Beyond Correlation: A Path-Invariant Measure for Seismogram Similarity
- Dual-Modal Attention-Enhanced Text-Video Retrieval with Triplet Partial Margin Contrastive Learning
- Online Self-Supervised Learning for Object Picking: Detecting Optimum Grasping Position using a Metric Learning Approach
- Generalized Inter-class Loss for Gait Recognition
- Cyber Vaccine for Deepfake Immunity
- Writer Identification and Writer Retrieval Based on NetVLAD with Re-ranking
- GANravel: User-Driven Direction Disentanglement in Generative Adversarial Networks
- Generative One-Shot Learning (GOL): A Semi-Parametric Approach to One-Shot Learning in Autonomous Vision
- Auxiliary Cross-Modal Representation Learning with Triplet Loss Functions for Online Handwriting Recognition
- Open-set Face Recognition for Small Galleries Using Siamese Networks
- Person Recognition in Personal Photo Collections
- Compositional Clustering: Applications to Multi-Label Object Recognition and Speaker Identification
- Intra-Variable Handwriting Inspection Reinforced with Idiosyncrasy Analysis
- The Impact of Racial Distribution in Training Data on Face Recognition Bias: A Closer Look
- Review of Demographic Fairness in Face Recognition
- Adversarial Margin Maximization Networks
- Towards a General Deep Feature Extractor for Facial Expression Recognition
- GANprintR: Improved Fakes and Evaluation of the State of the Art in Face Manipulation Detection
- Octuplet Loss: Make Face Recognition Robust to Image Resolution
- Few-Shot Learning with Uncertainty-based Quadruplet Selection for Interference Classification in GNSS Data
- Kinship Verification Based on Cross-Generation Feature Interaction Learning
- OSDFace: One-Step Diffusion Model for Face Restoration
- Novelty Detection and Analysis of Traffic Scenario Infrastructures in the Latent Space of a Vision Transformer-Based Triplet Autoencoder
- Shared Manifold Learning Using a Triplet Network for Multiple Sensor Translation and Fusion with Missing Data
- Consistent and Flexible Selectivity Estimation for High-Dimensional Data
- Semi-supervised Network Embedding with Differentiable Deep Quantisation
- LIDSNet: A Lightweight on-device Intent Detection model using Deep Siamese Network
- Learnable Optimal Sequential Grouping for Video Scene Detection
- Learning to rank music tracks using triplet loss
- Expert-LaSTS: Expert-Knowledge Guided Latent Space for Traffic Scenarios
- Spatio-Temporal Fusion Networks for Action Recognition
- A Light Heterogeneous Graph Collaborative Filtering Model using Textual Information
- HeartSiam: A Domain Invariant Model for Heart Sound Classification
- Personalizing Task-oriented Dialog Systems via Zero-shot Generalizable Reward Function
- A Fair Experimental Comparison of Neural Network Architectures for Latent Representations of Multi-Omics for Drug Response Prediction
- Towards minimizing efforts for Morphing Attacks -- Deep embeddings for morphing pair selection and improved Morphing Attack Detection
- ToMoBrush: Exploring Dental Health Sensing using a Sonic Toothbrush
- Building Computationally Efficient and Well-Generalizing Person Re-Identification Models with Metric Learning
- DeConFuse : A Deep Convolutional Transform based Unsupervised Fusion Framework
- Connecting Dualities and Machine Learning
- Enhancing Robustness of On-line Learning Models on Highly Noisy Data
- Semantic Cluster Unary Loss for Efficient Deep Hashing
- FingerSlid: Towards Finger-Sliding Continuous Authentication on Smart Devices Via Vibration
- openFEAT: Improving Speaker Identification by Open-set Few-shot Embedding Adaptation with Transformer
- From Face Recognition to Models of Identity: A Bayesian Approach to Learning about Unknown Identities from Unsupervised Data
- A False Sense of Privacy: Towards a Reliable Evaluation Methodology for the Anonymization of Biometric Data
- Learning Music-Dance Representations through Explicit-Implicit Rhythm Synchronization
- Contrastive-mixup learning for improved speaker verification
- VISHIEN-MAAT: Scrollytelling visualization design for explaining Siamese Neural Network concept to non-technical users
- PhishGAN: Data Augmentation and Identification of Homoglpyh Attacks
- Estimating and abstracting the 3D structure of bones using neural networks on X-ray (2D) images
- An End-to-End Network for Co-Saliency Detection in One Single Image
- Compact Network Training for Person ReID
- Semi-Supervised Learning using Siamese Networks
- Confidence-Calibrated Face and Kinship Verification
- Non-contrastive representation learning for intervals from well logs
- Supervision and Source Domain Impact on Representation Learning: A Histopathology Case Study
- A Temporal Sequence Learning for Action Recognition and Prediction
- Semantically Invariant Text-to-Image Generation
- Learning Effective Embeddings From Crowdsourced Labels: An Educational Case Study
- DimCL: Dimensional Contrastive Learning For Improving Self-Supervised Learning
- Model Guided Road Intersection Classification
- Rethinking Impersonation and Dodging Attacks on Face Recognition Systems
- RaLF: Flow-based Global and Metric Radar Localization in LiDAR Maps
- Efficient Large-Scale Face Clustering Using an Online Mixture of Gaussians
- ActiveGuard: An Active DNN IP Protection Technique via Adversarial Examples
- Improving safety in physical human-robot collaboration via deep metric learning
- A Novel Hybrid Scheme Using Genetic Algorithms and Deep Learning for the Reconstruction of Portuguese Tile Panels
- Contrastive Conditional Latent Diffusion for Audio-visual Segmentation
- Productive Crop Field Detection: A New Dataset and Deep Learning Benchmark Results
- DeePLT: Personalized Lighting Facilitates by Trajectory Prediction of Recognized Residents in the Smart Home
- Magnification Generalization for Histopathology Image Embedding
- Enhancing Sample Utilization in Noise-Robust Deep Metric Learning With Subgroup-Based Positive-Pair Selection
- ContrastNER: Contrastive-based Prompt Tuning for Few-shot NER
- Learning icons appearance similarity
- Proactive Prioritization of App Issues via Contrastive Learning
- Taking Modality-free Human Identification as Zero-shot Learning
- Ulixes: Facial Recognition Privacy with Adversarial Machine Learning
- Animating Through Warping: an Efficient Method for High-Quality Facial Expression Animation
- Recursive Chaining of Reversible Image-to-image Translators For Face Aging
- A Comparison of Metric Learning Loss Functions for End-To-End Speaker Verification
- 6MapNet: Representing soccer players from tracking data by a triplet network
- YOLO11-JDE: Fast and Accurate Multi-Object Tracking with Self-Supervised Re-ID
- FINER: Enhancing State-of-the-art Classifiers with Feature Attribution to Facilitate Security Analysis
- Universal Adversarial Spoofing Attacks against Face Recognition
- Individual Identification Using Radar-Measured Respiratory and Heartbeat Features
- Remote Sensor Design for Visual Recognition with Convolutional Neural Networks
- A Novel Semisupervised Contrastive Regression Framework for Forest Inventory Mapping with Multisensor Satellite Data
- A Survey on Personalized Content Synthesis with Diffusion Models
- Unsupervised Domain Adaptation for Cross-sensor Pore Detection in High-resolution Fingerprint Images
- Adaptive Affinity for Associations in Multi-Target Multi-Camera Tracking
- Recovering Faces from Portraits with Auxiliary Facial Attributes
- BiometricBlender: Ultra-high dimensional, multi-class synthetic data generator to imitate biometric feature space
- Understanding Hyperbolic Metric Learning through Hard Negative Sampling
- Semi-supervised Adversarial Learning for Complementary Item Recommendation
- iCub Knows Where You Look: Exploiting Social Cues for Interactive Object Detection Learning
- Transform, Contrast and Tell: Coherent Entity-Aware Multi-Image Captioning
- Context-Enriched Contrastive Loss: Enhancing Presentation of Inherent Sample Connections in Contrastive Learning Framework
- Machine-learned models for magnetic materials
- A Temporal-Spectral Fusion Transformer with Subject-Specific Adapter for Enhancing RSVP-BCI Decoding
- A Survey on Self-Supervised Graph Foundation Models: Knowledge-Based Perspective
- Autonomous Learning for Face Recognition in the Wild via Ambient Wireless Cues
- Biometric Authentication Based on Enhanced Remote Photoplethysmography Signal Morphology
- EfficientWord-Net: An Open Source Hotword Detection Engine based on One-shot Learning
- Learning Deep Convolutional Embeddings for Face Representation Using Joint Sample- and Set-based Supervision
- Overcoming Deceptiveness in Fitness Optimization with Unsupervised Quality-Diversity
- Cross-Global Attention Graph Kernel Network Prediction of Drug Prescription
- Geometrically Mappable Image Features
- Image-Text Pre-Training for Logo Recognition
- Discriminatory and orthogonal feature learning for noise robust keyword spotting
- Steganography Beyond Space-Time with Chain of Multimodal AI
- Artificial Immune System of Secure Face Recognition Against Adversarial Attacks
- Learning Test-time Augmentation for Content-based Image Retrieval
- CROLoss: Towards a Customizable Loss for Retrieval Models in Recommender Systems
- Atom Cloud Detection Using a Deep Neural Network
- Prototype Memory for Large-scale Face Representation Learning
- Rotation-Adaptive Point Cloud Domain Generalization via Intricate Orientation Learning
- Open-Set Biometrics: Beyond Good Closed-Set Models
- Distance-learning For Approximate Bayesian Computation To Model a Volcanic Eruption
- Whose Emotion Matters? Speaking Activity Localisation without Prior Knowledge
- The Importance of Context When Recommending TV Content: Dataset and Algorithms
- A non-discriminatory approach to ethical deep learning
- Robust Character Labeling in Movie Videos: Data Resources and Self-supervised Feature Adaptation
- A Cluster-Matching-Based Method for Video Face Recognition
- Music Recommendations in Hyperbolic Space: An Application of Empirical Bayes and Hierarchical Poincaré Embeddings
- Relevance-based Margin for Contrastively-trained Video Retrieval Models
- Deep Metric Learning with Alternating Projections onto Feasible Sets
- Part-based Multi-stream Model for Vehicle Searching
- Representation Learning for Tablet and Paper Domain Adaptation in Favor of Online Handwriting Recognition
- Transfer Learning and Augmentation for Word Sense Disambiguation
- re-OBJ: Jointly Learning the Foreground and Background for Object Instance Re-identification
- Using a GAN to Generate Adversarial Examples to Facial Image Recognition
- A Review on Self-Supervised Learning for Time Series Anomaly Detection: Recent Advances and Open Challenges
- From Multimodal to Unimodal Attention in Transformers using Knowledge Distillation
- CL2R: Compatible Lifelong Learning Representations
- A Large-Scale Re-identification Analysis in Sporting Scenarios: the Betrayal of Reaching a Critical Point
- Face Attributes as Cues for Deep Face Recognition Understanding
- Enhancing Function Name Prediction using Votes-Based Name Tokenization and Multi-Task Learning
- SynMorph: Generating Synthetic Face Morphing Dataset with Mated Samples
- CoReS: Compatible Representations via Stationarity
- Transfer of Pretrained Model Weights Substantially Improves Semi-Supervised Image Classification
- FACT: Foundation Model for Assessing Cancer Tissue Margins with Mass Spectrometry
- VIPeR: Visual Incremental Place Recognition with Adaptive Mining and Continual Learning
- From Detection to Action Recognition: An Edge-Based Pipeline for Robot Human Perception
- Semantic Similarity Measure of Natural Language Text through Machine Learning and a Keyword-Aware Cross-Encoder-Ranking Summarizer -- A Case Study Using UCGIS GIS&T Body of Knowledge
- Weaponizing Unicodes with Deep Learning -- Identifying Homoglyphs with Weakly Labeled Data
- Perona: Robust Infrastructure Fingerprinting for Resource-Efficient Big Data Analytics
- Learning Image Representations for Content Based Image Retrieval of Radiotherapy Treatment Plans
- TypoSwype: An Imaging Approach to Detect Typo-Squatting
- ElasticHash: Semantic Image Similarity Search by Deep Hashing with Elasticsearch
- Unsupervised Visual Time-Series Representation Learning and Clustering
- Universal Embedding Function for Traffic Classification via QUIC Domain Recognition Pretraining: A Transfer Learning Success
- Bilateral Asymmetry Guided Counterfactual Generating Network for Mammogram Classification
- An Efficient Method for Face Quality Assessment on the Edge
- Angular Triplet Loss-based Camera Network for ReID
- Joint Learning of Generative Translator and Classifier for Visually Similar Classes
- DISCERN: Diversity-based Selection of Centroids for k-Estimation and Rapid Non-stochastic Clustering
- Re-pseudonymization Strategies for Smart Meter Data Are Not Robust to Deep Learning Profiling Attacks
- I2CKD : Intra- and Inter-Class Knowledge Distillation for Semantic Segmentation
- Supervised Training of Siamese Spiking Neural Networks with Earth Mover's Distance
- Multi-modal data generation with a deep metric variational autoencoder
- User profiles matching for different social networks based on faces embeddings
- Classification and Clustering of Sentence-Level Embeddings of Scientific Articles Generated by Contrastive Learning
- My Eyes Are Up Here: Promoting Focus on Uncovered Regions in Masked Face Recognition
- Visual Persuasion in COVID-19 Social Media Content: A Multi-Modal Characterization
- A semantic embedding space based on large language models for modelling human beliefs
- DeepDiffusion: Unsupervised Learning of Retrieval-adapted Representations via Diffusion-based Ranking on Latent Feature Manifold
- Supervised Representation Learning towards Generalizable Assembly State Recognition
- Entropy-regularized Optimal Transport Generative Models
- Interpreting Face Inference Models using Hierarchical Network Dissection
- Finding Person Relations in Image Data of the Internet Archive
- Mitosis Detection Under Limited Annotation: A Joint Learning Approach
- Understanding and Exploiting Dependent Variables with Deep Metric Learning
- Learnable Adaptive Cosine Estimator (LACE) for Image Classification
- Nowhere to Hide: Cross-modal Identity Leakage between Biometrics and Devices
- Batch-Incremental Triplet Sampling for Training Triplet Networks Using Bayesian Updating Theorem
- Image Similarity using An Ensemble of Context-Sensitive Models
- Joint Discriminative and Metric Embedding Learning for Person Re-Identification
- Generalized Contrastive Learning for Multi-Modal Retrieval and Ranking
- ENCODE: Breaking the Trade-Off Between Performance and Efficiency in Long-Term User Behavior Modeling
- Are Adaptive Face Recognition Systems still Necessary? Experiments on the APE Dataset
- Self-supervised Latent Space Optimization with Nebula Variational Coding
- Learning Representations for Masked Facial Recovery
- Triplet loss based embeddings for forensic speaker identification in Spanish
- A Strong Baseline for Fashion Retrieval with Person Re-Identification Models
- A Set of Distinct Facial Traits Learned by Machines Is Not Predictive of Appearance Bias in the Wild
- Recommending Multiple Positive Citations for Manuscript via Content-Dependent Modeling and Multi-Positive Triplet
- Iterative Window Mean Filter: Thwarting Diffusion-based Adversarial Purification
- SwipeGANSpace: Swipe-to-Compare Image Generation via Efficient Latent Space Exploration
- Neural Network-based Information Set Weighting for Playing Reconnaissance Blind Chess
- Focal Inferential Infusion Coupled with Tractable Density Discrimination for Implicit Hate Detection
- Towards Better Modeling with Missing Data: A Contrastive Learning-based Visual Analytics Perspective
- Revisiting Experience Replayable Conditions
- Generating 2D and 3D Master Faces for Dictionary Attacks with a Network-Assisted Latent Space Evolution
- Informative Sample-Aware Proxy for Deep Metric Learning
- A Novel Graph-Theoretic Deep Representation Learning Method for Multi-Label Remote Sensing Image Retrieval
- Leveraging edge detection and neural networks for better UAV localization
- Modelling antimicrobial prescriptions in Scotland: A spatio-temporal clustering approach
- Learning Effective Representations for Retrieval Using Self-Distillation with Adaptive Relevance Margins
- Trajectory saliency detection using consistency-oriented latent codes from a recurrent auto-encoder
- CLOP: Video-and-Language Pre-Training with Knowledge Regularizations
- Visual-Informed Speech Enhancement Using Attention-Based Beamforming
- EfficientRec an unlimited user-item scale recommendation system based on clustering and users interaction embedding profile
- GE2E-KWS: Generalized End-to-End Training and Evaluation for Zero-shot Keyword Spotting
- A Clustering-Based Method for Automatic Educational Video Recommendation Using Deep Face-Features of Lecturers
- TripletViNet: Mitigating Misinformation Video Spread Across Platforms
- Knowledge Elicitation using Deep Metric Learning and Psychometric Testing
- MOFHEI: Model Optimizing Framework for Fast and Efficient Homomorphically Encrypted Neural Network Inference
- Contrastive Learning-Enhanced Large Language Models for Monolith-to-Microservice Decomposition
- REACT: Real-time Efficient Attribute Clustering and Transfer for Updatable 3D Scene Graph
- Development of a defacing algorithm to protect the privacy of head and neck cancer patients in publicly-accessible radiotherapy datasets
- A prototype neutron-detector array for future deep-underground s-process studies
- Privacy-Preserving Person Re-Identification from Temporal Sequences with Transformer and Hungarian Optimization
- Adaptive Discriminative Regularization for Visual Classification
- Set-Based Face Recognition Beyond Disentanglement: Burstiness Suppression With Variance Vocabulary
- Contrastive Knowledge Distillation for Anomaly Detection in Multi-Illumination/Focus Display Images
- WildGait: Learning Gait Representations from Raw Surveillance Streams
- Learning Continuous Face Representation with Explicit Functions
- Open-Set Face Identification on Few-Shot Gallery by Fine-Tuning
- DCSCR: A Class-Specific Collaborative Representation based Network for Image Set Classification
- Além do Desempenho: Um Estudo da Confiabilidade de Detectores de Deepfakes
- PD-Loss: Proxy-Decidability for Efficient Metric Learning
- Informative and Representative Triplet Selection for Multilabel Remote Sensing Image Retrieval
- Lost in Context: The Influence of Context on Feature Attribution Methods for Object Recognition
- Few-Shot Network Intrusion Detection Using Online Triplet Mining
- A Novel Method for News Article Event-Based Embedding
- SIG: A Synthetic Identity Generation Pipeline for Generating Evaluation Datasets for Face Recognition
- Identifying Scientists on X
- Face and Voice Cross-modal Association with Learning Convex Feature Embedding
- Perspective-Invariant Attack with Enhanced Transferability of Adversarial Examples
- A Framework for Mining Collectively-Behaving Bots in MMORPGs
- Scalable Deep Metric Learning on Attributed Graphs
- Boosting Network Weight Separability via Feed-Backward Reconstruction
- Multi-Grained Vision-Language Alignment for Domain Generalized Person Re-Identification
- Unveiling the Potential: Harnessing Deep Metric Learning to Circumvent Video Streaming Encryption
- Supervised Fine-tuning Evaluation for Long-term Visual Place Recognition
- FasterVideo: Efficient Online Joint Object Detection And Tracking
- Surrogate assisted diversity estimation in neural ensemble search
- Revealing Unintentional Information Leakage in Low-Dimensional Facial Portrait Representations
- Identifying partial mouse brain microscopy images from Allen reference atlas using a contrastively learned semantic space
- Improving Generalization in Deepfake Detection with Face Foundation Models and Metric Learning
- An Empirical Study of Visual Features for DNN based Audio-Visual Speech Enhancement in Multi-talker Environments
- ExeChecker: Where Did I Go Wrong?
- That was not what I was aiming at! Differentiating human intent and outcome in a physically dynamic throwing task