Very Deep Convolutional Networks for Large-Scale Image Recognition
arXiv:1409.1556
Abstract
In this work we investigate the effect of the convolutional network depth on its accuracy in the large-scale image recognition setting. Our main contribution is a thorough evaluation of networks of increasing depth using an architecture with very small (3x3) convolution filters, which shows that a significant improvement on the prior-art configurations can be achieved by pushing the depth to 16-19 weight layers. These findings were the basis of our ImageNet Challenge 2014 submission, where our team secured the first and the second places in the localisation and classification tracks respectively. We also show that our representations generalise well to other datasets, where they achieve state-of-the-art results. We have made our two best-performing ConvNet models publicly available to facilitate further research on the use of deep visual representations in computer vision.
References in corpus (7)
- Going Deeper with Convolutions
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- One weird trick for parallelizing convolutional neural networks
- Deep Visual-Semantic Alignments for Generating Image Descriptions
- Deep convolutional filter banks for texture recognition and segmentation
- Material Recognition in the Wild with the Materials in Context Database
- Actions and Attributes from Wholes and Parts
Cited by in corpus (2013)
- Neural Architecture Search with Reinforcement Learning
- Convolutional Neural Networks for Medical Image Analysis: Full Training or Fine Tuning?
- Striving for Simplicity: The All Convolutional Net
- PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space
- FitNets: Hints for Thin Deep Nets
- Beyond Correlation Filters: Learning Continuous Convolution Operators for Visual Tracking
- NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding
- Learning Face Representation from Scratch
- UnitBox: An Advanced Object Detection Network
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- FINN: A Framework for Fast, Scalable Binarized Neural Network Inference
- Person Re-identification: Past, Present and Future
- Compressing Deep Convolutional Networks using Vector Quantization
- DeepVO: Towards End-to-End Visual Odometry with Deep Recurrent Convolutional Neural Networks
- Blind Image Quality Assessment Using A Deep Bilinear Convolutional Neural Network
- DeepID3: Face Recognition with Very Deep Neural Networks
- ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices
- Applications of Deep Learning and Reinforcement Learning to Biological Data
- REFUGE Challenge: A Unified Framework for Evaluating Automated Methods for Glaucoma Assessment from Fundus Photographs
- Network Trimming: A Data-Driven Neuron Pruning Approach towards Efficient Deep Architectures
- Cost-Effective Active Learning for Deep Image Classification
- PlantDoc: A Dataset for Visual Plant Disease Detection
- Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
- A Survey of Recent Advances in CNN-based Single Image Crowd Counting and Density Estimation
- HDR image reconstruction from a single exposure using deep CNNs
- TernausNet: U-Net with VGG11 Encoder Pre-Trained on ImageNet for Image Segmentation
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- Grid Search, Random Search, Genetic Algorithm: A Big Comparison for NAS
- Training Convolutional Networks with Noisy Labels
- FPGA-based Accelerators of Deep Learning Networks for Learning and Classification: A Review
- R2CNN: Rotational Region CNN for Orientation Robust Scene Text Detection
- Learning Spatial Fusion for Single-Shot Object Detection
- Learning Structured Sparsity in Deep Neural Networks
- Adversarial Discriminative Domain Adaptation
- TextBoxes: A Fast Text Detector with a Single Deep Neural Network
- YOLO9000: Better, Faster, Stronger
- Visualizing Deep Neural Network Decisions: Prediction Difference Analysis
- GLAD: Global-Local-Alignment Descriptor for Pedestrian Retrieval
- Visual Saliency Detection Based on Multiscale Deep CNN Features
- Residual Networks of Residual Networks: Multilevel Residual Networks
- An Empirical Evaluation of Deep Learning on Highway Driving
- Massively Parallel Methods for Deep Reinforcement Learning
- Do ImageNet Classifiers Generalize to ImageNet?
- Cascade R-CNN: Delving into High Quality Object Detection
- Advancements in Image Classification using Convolutional Neural Network
- Learning Features for Offline Handwritten Signature Verification using Deep Convolutional Neural Networks
- An In-field Automatic Wheat Disease Diagnosis System
- Explain Images with Multimodal Recurrent Neural Networks
- DeepNAT: Deep Convolutional Neural Network for Segmenting Neuroanatomy
- Fast-SCNN: Fast Semantic Segmentation Network
- CSPNet: A New Backbone that can Enhance Learning Capability of CNN
- Channel Pruning for Accelerating Very Deep Neural Networks
- Person Re-Identification by Camera Correlation Aware Feature Augmentation
- Speeding-up Convolutional Neural Networks Using Fine-tuned CP-Decomposition
- Deep Image: Scaling up Image Recognition
- Deep Learning of Subsurface Flow via Theory-guided Neural Network
- VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection
- Standardized Assessment of Automatic Segmentation of White Matter Hyperintensities and Results of the WMH Segmentation Challenge
- What makes ImageNet good for transfer learning?
- Retinal vessel segmentation based on Fully Convolutional Neural Networks
- Revisiting Unreasonable Effectiveness of Data in Deep Learning Era
- Deep learning electromagnetic inversion with convolutional neural networks
- Multi-source Transfer Learning with Convolutional Neural Networks for Lung Pattern Analysis
- Exploring Generalization in Deep Learning
- Domain Adaptation for Visual Applications: A Comprehensive Survey
- Person Re-identification by Contour Sketch under Moderate Clothing Change
- Monte-Carlo Sampling applied to Multiple Instance Learning for Histological Image Classification
- Fully Convolutional Multi-Class Multiple Instance Learning
- Fully Connected Deep Structured Networks
- Variational Approaches for Auto-Encoding Generative Adversarial Networks
- Deep Learning for Human Affect Recognition: Insights and New Developments
- Embedding Structured Contour and Location Prior in Siamesed Fully Convolutional Networks for Road Detection
- Diagnose like a Radiologist: Attention Guided Convolutional Neural Network for Thorax Disease Classification
- SoundNet: Learning Sound Representations from Unlabeled Video
- Deep Gaze I: Boosting Saliency Prediction with Feature Maps Trained on ImageNet
- Exploring the Regularity of Sparse Structure in Convolutional Neural Networks
- Wider or Deeper: Revisiting the ResNet Model for Visual Recognition
- Weakly Supervised Adversarial Domain Adaptation for Semantic Segmentation in Urban Scenes
- Combining Physically-Based Modeling and Deep Learning for Fusing GRACE Satellite Data: Can We Learn from Mismatch?
- DeepSign: Deep Learning for Automatic Malware Signature Generation and Classification
- Fast Patch-based Style Transfer of Arbitrary Style
- Neural Paraphrase Generation with Stacked Residual LSTM Networks
- Joint Activity Recognition and Indoor Localization with WiFi Fingerprints
- Star-galaxy Classification Using Deep Convolutional Neural Networks
- Deep Transfer Learning for Person Re-identification
- Spectral Norm Regularization for Improving the Generalizability of Deep Learning
- GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond
- Detecting Curve Text in the Wild: New Dataset and New Solution
- Scene Text Detection via Holistic, Multi-Channel Prediction
- Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network
- Few-Shot Adversarial Domain Adaptation
- Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World
- Driver Drowsiness Detection Model Using Convolutional Neural Networks Techniques for Android Application
- Exploiting Unlabeled Data in CNNs by Self-supervised Learning to Rank
- Learning to Compare Image Patches via Convolutional Neural Networks
- Deep Adaptive Feature Embedding with Local Sample Distributions for Person Re-identification
- MoCap-guided Data Augmentation for 3D Pose Estimation in the Wild
- Spatially-sparse convolutional neural networks
- The Reversible Residual Network: Backpropagation Without Storing Activations
- Fathom: Reference Workloads for Modern Deep Learning Methods
- A survey of advances in vision-based vehicle re-identification
- Spatial Group-wise Enhance: Improving Semantic Feature Learning in Convolutional Networks
- CirCNN: Accelerating and Compressing Deep Neural Networks Using Block-CirculantWeight Matrices
- SuperNeurons: Dynamic GPU Memory Management for Training Deep Neural Networks
- Ablation Studies in Artificial Neural Networks
- Places: An Image Database for Deep Scene Understanding
- Gazelle: A Low Latency Framework for Secure Neural Network Inference
- Understanding urban landuse from the above and ground perspectives: a deep learning, multimodal solution
- FusionNet: 3D Object Classification Using Multiple Data Representations
- Learning Visual Importance for Graphic Designs and Data Visualizations
- Training Deeper Convolutional Networks with Deep Supervision
- Simple Black-Box Adversarial Perturbations for Deep Networks
- Deeply-Learned Part-Aligned Representations for Person Re-Identification
- Ensemble of Deep Convolutional Neural Networks for Automatic Pavement Crack Detection and Measurement
- JALAD: Joint Accuracy- and Latency-Aware Deep Structure Decoupling for Edge-Cloud Execution
- Recent Advance in Content-based Image Retrieval: A Literature Survey
- Exploring Nearest Neighbor Approaches for Image Captioning
- Visual Entailment: A Novel Task for Fine-Grained Image Understanding
- A Comprehensive guide to Bayesian Convolutional Neural Network with Variational Inference
- Less-forgetting Learning in Deep Neural Networks
- From Google Maps to a Fine-Grained Catalog of Street trees
- End-to-End Photo-Sketch Generation via Fully Convolutional Representation Learning
- A Joint Convolutional Neural Networks and Context Transfer for Street Scenes Labeling
- Context-Aware Visual Policy Network for Fine-Grained Image Captioning
- What-and-Where to Match: Deep Spatially Multiplicative Integration Networks for Person Re-identification
- Improving Neural Network Quantization without Retraining using Outlier Channel Splitting
- Deep Visual-Semantic Alignments for Generating Image Descriptions
- Bag of Freebies for Training Object Detection Neural Networks
- An Entropy-based Pruning Method for CNN Compression
- AutoSlim: Towards One-Shot Architecture Search for Channel Numbers
- DSD: Dense-Sparse-Dense Training for Deep Neural Networks
- Palmprint Recognition in Uncontrolled and Uncooperative Environment
- Dynamic Network Surgery for Efficient DNNs
- Deep EndoVO: A Recurrent Convolutional Neural Network (RCNN) based Visual Odometry Approach for Endoscopic Capsule Robots
- ImageNet pre-trained models with batch normalization
- Audio Spectrogram Representations for Processing with Convolutional Neural Networks
- A Boundary Tilting Persepective on the Phenomenon of Adversarial Examples
- Class-Balanced Loss Based on Effective Number of Samples
- Quantifying the Carbon Emissions of Machine Learning
- Deep Convolutions for In-Depth Automated Rock Typing
- Deep learning for source camera identification on mobile devices
- The Expressive Power of Neural Networks: A View from the Width
- Drivers Drowsiness Detection using Condition-Adaptive Representation Learning Framework
- An Implementation of Faster RCNN with Study for Region Sampling
- Deep Learning: A Bayesian Perspective
- Men Also Like Shopping: Reducing Gender Bias Amplification using Corpus-level Constraints
- Tropical Cyclone Track Forecasting using Fused Deep Learning from Aligned Reanalysis Data
- CrowdNet: A Deep Convolutional Network for Dense Crowd Counting
- Robustness of classifiers: from adversarial to random noise
- A Bayesian Data Augmentation Approach for Learning Deep Models
- XNOR-Net++: Improved Binary Neural Networks
- N2N Learning: Network to Network Compression via Policy Gradient Reinforcement Learning
- Amulet: Aggregating Multi-level Convolutional Features for Salient Object Detection
- VoxResNet: Deep Voxelwise Residual Networks for Volumetric Brain Segmentation
- Gated-SCNN: Gated Shape CNNs for Semantic Segmentation
- Evaluating Two-Stream CNN for Video Classification
- TimeNet: Pre-trained deep recurrent neural network for time series classification
- Single-Shot Refinement Neural Network for Object Detection
- Visual Relationship Detection with Language Priors
- SaltiNet: Scan-path Prediction on 360 Degree Images using Saliency Volumes
- Discover Your Social Identity from What You Tweet: a Content Based Approach
- Vista: A Visually, Socially, and Temporally-aware Model for Artistic Recommendation
- B-CNN: Branch Convolutional Neural Network for Hierarchical Classification
- Deeper and Wider Siamese Networks for Real-Time Visual Tracking
- Discriminative Active Learning
- Rethinking on Multi-Stage Networks for Human Pose Estimation
- Look Wider to Match Image Patches with Convolutional Neural Networks
- Root Mean Square Layer Normalization
- Scale-Invariant Convolutional Neural Networks
- Fast Feature Fool: A data independent approach to universal adversarial perturbations
- ThiNet: A Filter Level Pruning Method for Deep Neural Network Compression
- Learning Deep Structured Multi-Scale Features using Attention-Gated CRFs for Contour Prediction
- Federated Learning with Matched Averaging
- Edge Intelligence: Paving the Last Mile of Artificial Intelligence with Edge Computing
- Wavelet Convolutional Neural Networks for Texture Classification
- Deep Model Compression: Distilling Knowledge from Noisy Teachers
- Switching Convolutional Neural Network for Crowd Counting
- InstaGAN: Instance-aware Image-to-Image Translation
- Approaching the Computational Color Constancy as a Classification Problem through Deep Learning
- A Deep Learning Perspective on the Origin of Facial Expressions
- PixelNet: Representation of the pixels, by the pixels, and for the pixels
- Move Evaluation in Go Using Deep Convolutional Neural Networks
- Re-ID done right: towards good practices for person re-identification
- Universal adversarial perturbations
- Classification of Time-Series Images Using Deep Convolutional Neural Networks
- Perceptual Generative Adversarial Networks for Small Object Detection
- Generalized Inner Loop Meta-Learning
- RECOD Titans at ISIC Challenge 2017
- On the Origin of Deep Learning
- Deep Active Learning over the Long Tail
- Deep learning for galaxy surface brightness profile fitting
- Hierarchical Object Detection with Deep Reinforcement Learning
- Bi-directional Dermoscopic Feature Learning and Multi-scale Consistent Decision Fusion for Skin Lesion Segmentation
- A novel database of Children's Spontaneous Facial Expressions (LIRIS-CSE)
- Tree-Structured Reinforcement Learning for Sequential Object Localization
- Deep Convolutional Neural Network Inference with Floating-point Weights and Fixed-point Activations
- Best of Both Worlds: Transferring Knowledge from Discriminative Learning to a Generative Visual Dialog Model
- Exploiting Local Features from Deep Networks for Image Retrieval
- Peephole: Predicting Network Performance Before Training
- Automatic Liver Lesion Detection using Cascaded Deep Residual Networks
- Fisher Vectors Derived from Hybrid Gaussian-Laplacian Mixture Models for Image Annotation
- Defeating Image Obfuscation with Deep Learning
- LiteSeg: A Novel Lightweight ConvNet for Semantic Segmentation
- Learning Robust Representations via Multi-View Information Bottleneck
- Clipper: A Low-Latency Online Prediction Serving System
- PVANet: Lightweight Deep Neural Networks for Real-time Object Detection
- Angle-Closure Detection in Anterior Segment OCT based on Multi-Level Deep Network
- Seizure Detection using Least EEG Channels by Deep Convolutional Neural Network
- Visual Madlibs: Fill in the blank Image Generation and Question Answering
- Dissecting the Graphcore IPU Architecture via Microbenchmarking
- Interleaved Group Convolutions for Deep Neural Networks
- Detecting motorcycle helmet use with deep learning
- DeepID-Net: Deformable Deep Convolutional Neural Networks for Object Detection
- Exploiting Image-trained CNN Architectures for Unconstrained Video Classification
- Real Time Image Saliency for Black Box Classifiers
- Data Augmentation in Emotion Classification Using Generative Adversarial Networks
- POBA-GA: Perturbation Optimized Black-Box Adversarial Attacks via Genetic Algorithm
- Practical Approaches Towards Deep-Learning Based Cross-Device Power Side Channel Attack
- Local Aggregation for Unsupervised Learning of Visual Embeddings
- NeuralPower: Predict and Deploy Energy-Efficient Convolutional Neural Networks
- Hierarchical Attention Network for Action Recognition in Videos
- Recurrent Topic-Transition GAN for Visual Paragraph Generation
- Hierarchical Convolutional-Deconvolutional Neural Networks for Automatic Liver and Tumor Segmentation
- A Fully Convolutional Deep Auditory Model for Musical Chord Recognition
- UNet++: Redesigning Skip Connections to Exploit Multiscale Features in Image Segmentation
- Implicit Regularization in Deep Learning
- Non-rigid image registration using fully convolutional networks with deep self-supervision
- FreezeOut: Accelerate Training by Progressively Freezing Layers
- EmBench: Quantifying Performance Variations of Deep Neural Networks across Modern Commodity Devices
- Depth Adaptive Deep Neural Network for Semantic Segmentation
- StyleBank: An Explicit Representation for Neural Image Style Transfer
- Distributed Deep Neural Networks over the Cloud, the Edge and End Devices
- A Systematic Comparison of Bayesian Deep Learning Robustness in Diabetic Retinopathy Tasks
- Video Frame Interpolation via Adaptive Separable Convolution
- Towards seamless multi-view scene analysis from satellite to street-level
- Deep Convolution Networks for Compression Artifacts Reduction
- Do Convolutional Networks need to be Deep for Text Classification ?
- Scaling Deep Learning on GPU and Knights Landing clusters
- Fast YOLO: A Fast You Only Look Once System for Real-time Embedded Object Detection in Video
- Cyberbullying Detection in Social Networks Using Deep Learning Based Models; A Reproducibility Study
- Learning Background-Aware Correlation Filters for Visual Tracking
- ssEMnet: Serial-section Electron Microscopy Image Registration using a Spatial Transformer Network with Learned Features
- Deep Learning for Lung Cancer Detection: Tackling the Kaggle Data Science Bowl 2017 Challenge
- Improving Malaria Parasite Detection from Red Blood Cell using Deep Convolutional Neural Networks
- Illuminating Pedestrians via Simultaneous Detection & Segmentation
- DeepSZ: A Novel Framework to Compress Deep Neural Networks by Using Error-Bounded Lossy Compression
- Revisiting Self-Supervised Visual Representation Learning
- See, Hear, and Read: Deep Aligned Representations
- Understanding Top-k Sparsification in Distributed Deep Learning
- Detecting GAN-generated Imagery using Color Cues
- Deep Learning with Low Precision by Half-wave Gaussian Quantization
- Diverse and Accurate Image Description Using a Variational Auto-Encoder with an Additive Gaussian Encoding Space
- A Study of BFLOAT16 for Deep Learning Training
- Residual Convolutional CTC Networks for Automatic Speech Recognition
- Transformer-Transducer: End-to-End Speech Recognition with Self-Attention
- Co-training for Demographic Classification Using Deep Learning from Label Proportions
- How Can We Be So Dense? The Benefits of Using Highly Sparse Representations
- Expression Analysis Based on Face Regions in Read-world Conditions
- Salient Object Detection with Lossless Feature Reflection and Weighted Structural Loss
- Semi-Heterogeneous Three-Way Joint Embedding Network for Sketch-Based Image Retrieval
- FishNet: A Versatile Backbone for Image, Region, and Pixel Level Prediction
- Training Skinny Deep Neural Networks with Iterative Hard Thresholding Methods
- A Probabilistic Quality Representation Approach to Deep Blind Image Quality Prediction
- Improving Variational Autoencoder with Deep Feature Consistent and Generative Adversarial Training
- Sobolev Training for Neural Networks
- Feature Pyramid and Hierarchical Boosting Network for Pavement Crack Detection
- Learning Feature Pyramids for Human Pose Estimation
- Pyramid Feature Attention Network for Saliency detection
- Domain Adaptation for Object Detection via Style Consistency
- Machine learning for music genre: multifaceted review and experimentation with audioset
- Real-Time Illegal Parking Detection System Based on Deep Learning
- Real-time Convolutional Neural Networks for Emotion and Gender Classification
- Classification of Histopathological Biopsy Images Using Ensemble of Deep Learning Networks
- BranchyNet: Fast Inference via Early Exiting from Deep Neural Networks
- Cluster Pruning: An Efficient Filter Pruning Method for Edge AI Vision Applications
- Deep Matching Prior Network: Toward Tighter Multi-oriented Text Detection
- Analysis and Optimization of Convolutional Neural Network Architectures
- Deep Learning Based Large-Scale Automatic Satellite Crosswalk Classification
- Visual Relationship Detection with Internal and External Linguistic Knowledge Distillation
- -softmax: Improving Intra-class Compactness and Inter-class Separability of Features
- 2017 Robotic Instrument Segmentation Challenge
- Data Augmentation for Object Detection via Progressive and Selective Instance-Switching
- Parallel Tracking and Verifying: A Framework for Real-Time and High Accuracy Visual Tracking
- Context-Aware Embeddings for Automatic Art Analysis
- Deep Hyperspherical Learning
- Label Efficient Learning of Transferable Representations across Domains and Tasks
- PixelNet: Towards a General Pixel-level Architecture
- Learning Spatial Regularization with Image-level Supervisions for Multi-label Image Classification
- Modeling Context in Referring Expressions
- Online Multi-Object Tracking Using CNN-based Single Object Tracker with Spatial-Temporal Attention Mechanism
- Privacy Protection in Street-View Panoramas using Depth and Multi-View Imagery
- Revisiting the Effectiveness of Off-the-shelf Temporal Modeling Approaches for Large-scale Video Classification
- Word-level Deep Sign Language Recognition from Video: A New Large-scale Dataset and Methods Comparison
- Group Re-Identification with Multi-grained Matching and Integration
- PixelLink: Detecting Scene Text via Instance Segmentation
- Real-time Semantic Image Segmentation via Spatial Sparsity
- Multispectral Deep Neural Networks for Pedestrian Detection
- Multi-label Image Recognition by Recurrently Discovering Attentional Regions
- Deep Learning the City : Quantifying Urban Perception At A Global Scale
- PoseRBPF: A Rao-Blackwellized Particle Filter for 6D Object Pose Tracking
- segDeepM: Exploiting Segmentation and Context in Deep Neural Networks for Object Detection
- Zero-Shot Learning via Class-Conditioned Deep Generative Models
- OmniArt: Multi-task Deep Learning for Artistic Data Analysis
- TBC-Net: A real-time detector for infrared small target detection using semantic constraint
- RTM3D: Real-time Monocular 3D Detection from Object Keypoints for Autonomous Driving
- Learning from Synthetic Data for Crowd Counting in the Wild
- Deep learning for plasma tomography using the bolometer system at JET
- REMAP: Multi-layer entropy-guided pooling of dense CNN features for image retrieval
- Registration-free Face-SSD: Single shot analysis of smiles, facial attributes, and affect in the wild
- Range Loss for Deep Face Recognition with Long-tail
- Poverty Prediction with Public Landsat 7 Satellite Imagery and Machine Learning
- pCAMP: Performance Comparison of Machine Learning Packages on the Edges
- Black-Box Adversarial Attack with Transferable Model-based Embedding
- Accurate Pulmonary Nodule Detection in Computed Tomography Images Using Deep Convolutional Neural Networks
- Deep convolutional filter banks for texture recognition and segmentation
- Augment your batch: better training with larger batches
- Fast Scene Understanding for Autonomous Driving
- Optimizing and Visualizing Deep Learning for Benign/Malignant Classification in Breast Tumors
- Hermite-Gaussian Mode Detection via Convolution Neural Networks
- Poverty Mapping Using Convolutional Neural Networks Trained on High and Medium Resolution Satellite Images, With an Application in Mexico
- Relational Knowledge Distillation
- Biased Importance Sampling for Deep Neural Network Training
- Visual Causal Feature Learning
- Neural system identification for large populations separating "what" and "where"
- Design Space Exploration of Hardware Spiking Neurons for Embedded Artificial Intelligence
- Dense Prediction on Sequences with Time-Dilated Convolutions for Speech Recognition
- Interleaved Text/Image Deep Mining on a Large-Scale Radiology Database for Automated Image Interpretation
- Deep Convolutional Neural Network Design Patterns
- Image Captioning at Will: A Versatile Scheme for Effectively Injecting Sentiments into Image Descriptions
- YouTube-BoundingBoxes: A Large High-Precision Human-Annotated Data Set for Object Detection in Video
- Deep Architectures for Automated Seizure Detection in Scalp EEGs
- Dual Attention Networks for Multimodal Reasoning and Matching
- ACFNet: Attentional Class Feature Network for Semantic Segmentation
- Chart-Text: A Fully Automated Chart Image Descriptor
- Enabling FDD Massive MIMO through Deep Learning-based Channel Prediction
- CANet: Class-Agnostic Segmentation Networks with Iterative Refinement and Attentive Few-Shot Learning
- Deep Transfer Learning Methods for Colon Cancer Classification in Confocal Laser Microscopy Images
- Generalized BackPropagation, Étude De Cas: Orthogonality
- EvalAI: Towards Better Evaluation Systems for AI Agents
- WordSup: Exploiting Word Annotations for Character based Text Detection
- Visual Question Answering: A Survey of Methods and Datasets
- Sim2Real View Invariant Visual Servoing by Recurrent Control
- Low-memory GEMM-based convolution algorithms for deep neural networks
- Support Vector Guided Softmax Loss for Face Recognition
- Improving One-Shot Learning through Fusing Side Information
- Data Distillation: Towards Omni-Supervised Learning
- Evading Defenses to Transferable Adversarial Examples by Translation-Invariant Attacks
- Separating the EoR Signal with a Convolutional Denoising Autoencoder: A Deep-learning-based Method
- Boosting Occluded Image Classification via Subspace Decomposition Based Estimation of Deep Features
- Multi-Channel CNN-based Object Detection for Enhanced Situation Awareness
- Incorporating Network Built-in Priors in Weakly-supervised Semantic Segmentation
- Radar-based Feature Design and Multiclass Classification for Road User Recognition
- Deep Radiomics for Brain Tumor Detection and Classification from Multi-Sequence MRI
- StreetStyle: Exploring world-wide clothing styles from millions of photos
- BlackMarks: Blackbox Multibit Watermarking for Deep Neural Networks
- Recurrent Neural Network for (Un-)supervised Learning of Monocular VideoVisual Odometry and Depth
- Shift: A Zero FLOP, Zero Parameter Alternative to Spatial Convolutions
- Using Deep Learning to Localize Gravitational Wave Sources
- Deep Stacked Hierarchical Multi-patch Network for Image Deblurring
- Learning to Recognize 3D Human Action from A New Skeleton-based Representation Using Deep Convolutional Neural Networks
- Improved training of binary networks for human pose estimation and image recognition
- Two-Phase Learning for Weakly Supervised Object Localization
- Image Companding and Inverse Halftoning using Deep Convolutional Neural Networks
- Towards End-to-End Car License Plates Detection and Recognition with Deep Neural Networks
- Photo Aesthetics Ranking Network with Attributes and Content Adaptation
- Orthogonal Gradient Descent for Continual Learning
- Boosting Image Captioning with Attributes
- Untangling Local and Global Deformations in Deep Convolutional Networks for Image Classification and Sliding Window Detection
- When Face Recognition Meets with Deep Learning: an Evaluation of Convolutional Neural Networks for Face Recognition
- Condition-Invariant Multi-View Place Recognition
- Unsupervised End-to-end Learning for Deformable Medical Image Registration
- Graph-based Knowledge Distillation by Multi-head Attention Network
- FaceNet2ExpNet: Regularizing a Deep Face Recognition Net for Expression Recognition
- SPP-Net: Deep Absolute Pose Regression with Synthetic Views
- Learning A Physical Long-term Predictor
- Dimensionality Reduction using Similarity-induced Embeddings
- Trading-off Accuracy and Energy of Deep Inference on Embedded Systems: A Co-Design Approach
- Learning what to look in chest X-rays with a recurrent visual attention model
- Compact Deep Convolutional Neural Networks With Coarse Pruning
- Multimodal Attention for Neural Machine Translation
- DeepSpline: Data-Driven Reconstruction of Parametric Curves and Surfaces
- Mask-Guided Attention Network for Occluded Pedestrian Detection
- On the Robustness of Convolutional Neural Networks to Internal Architecture and Weight Perturbations
- Identity-Aware Textual-Visual Matching with Latent Co-attention
- Convolutional Neural Networks Analyzed via Convolutional Sparse Coding
- Stacked Conditional Generative Adversarial Networks for Jointly Learning Shadow Detection and Shadow Removal
- Instance-Aware Representation Learning and Association for Online Multi-Person Tracking
- EXTD: Extremely Tiny Face Detector via Iterative Filter Reuse
- GAP: Generalizable Approximate Graph Partitioning Framework
- Levelling the Playing Field: A Comprehensive Comparison of Visual Place Recognition Approaches under Changing Conditions
- 3G structure for image caption generation
- Generating Synthetic Data for Text Recognition
- Variational Information Distillation for Knowledge Transfer
- Heterogeneous Memory Enhanced Multimodal Attention Model for Video Question Answering
- C^3 Framework: An Open-source PyTorch Code for Crowd Counting
- SSA-CNN: Semantic Self-Attention CNN for Pedestrian Detection
- SC-DCNN: Highly-Scalable Deep Convolutional Neural Network using Stochastic Computing
- Acoustic scene classification using convolutional neural network and multiple-width frequency-delta data augmentation
- An attention-based multi-resolution model for prostate whole slide imageclassification and localization
- GFD-SSD: Gated Fusion Double SSD for Multispectral Pedestrian Detection
- Need for Speed: A Benchmark for Higher Frame Rate Object Tracking
- Interactive Video Object Segmentation in the Wild
- Can WiFi Estimate Person Pose?
- Sparsely-Connected Neural Networks: Towards Efficient VLSI Implementation of Deep Neural Networks
- Defensive Quantization: When Efficiency Meets Robustness
- Understanding and Comparing Deep Neural Networks for Age and Gender Classification
- Associatively Segmenting Instances and Semantics in Point Clouds
- Towards Efficient Training for Neural Network Quantization
- Multi-Modality Fusion based on Consensus-Voting and 3D Convolution for Isolated Gesture Recognition
- Style Transfer for Anime Sketches with Enhanced Residual U-net and Auxiliary Classifier GAN
- Quality Resilient Deep Neural Networks
- Deep Neural Network with l2-norm Unit for Brain Lesions Detection
- Real-time Convolutional Networks for Depth-based Human Pose Estimation
- Video Frame Interpolation via Adaptive Convolution
- SFD: Single Shot Scale-invariant Face Detector
- Curriculum Audiovisual Learning
- Robustness of Object Recognition under Extreme Occlusion in Humans and Computational Models
- Face Synthesis from Visual Attributes via Sketch using Conditional VAEs and GANs
- Concurrent Activity Recognition with Multimodal CNN-LSTM Structure
- Challenges in Disentangling Independent Factors of Variation
- Visual Attribute Transfer through Deep Image Analogy
- Doubly Convolutional Neural Networks
- Face Parsing via a Fully-Convolutional Continuous CRF Neural Network
- Hyper-Sphere Quantization: Communication-Efficient SGD for Federated Learning
- Compressing Convolutional Neural Networks
- DeepFL-IQA: Weak Supervision for Deep IQA Feature Learning
- Evaluating Saliency Map Explanations for Convolutional Neural Networks: A User Study
- Regularizing Deep Multi-Task Networks using Orthogonal Gradients
- Pruning Convolutional Neural Networks with Self-Supervision
- PipeCNN: An OpenCL-Based FPGA Accelerator for Large-Scale Convolution Neuron Networks
- Accurate and Compact Convolutional Neural Networks with Trained Binarization
- Unsupervised Degradation Learning for Single Image Super-Resolution
- Unsupervised Deep Tracking
- Clique pooling for graph classification
- BinaryDuo: Reducing Gradient Mismatch in Binary Activation Network by Coupling Binary Activations
- Low-light Image Enhancement Algorithm Based on Retinex and Generative Adversarial Network
- ViTOR: Learning to Rank Webpages Based on Visual Features
- Adapting Deep Network Features to Capture Psychological Representations
- A Computer Vision Pipeline for Automated Determination of Cardiac Structure and Function and Detection of Disease by Two-Dimensional Echocardiography
- Noisy Softmax: Improving the Generalization Ability of DCNN via Postponing the Early Softmax Saturation
- Ultrasound Image Representation Learning by Modeling Sonographer Visual Attention
- A Discriminative CNN Video Representation for Event Detection
- Visual pathways from the perspective of cost functions and multi-task deep neural networks
- Scene Graph Generation from Objects, Phrases and Region Captions
- Efficient Riemannian Optimization on the Stiefel Manifold via the Cayley Transform
- Jointly Modeling Embedding and Translation to Bridge Video and Language
- Graph-RISE: Graph-Regularized Image Semantic Embedding
- Style2Vec: Representation Learning for Fashion Items from Style Sets
- Built-in Foreground/Background Prior for Weakly-Supervised Semantic Segmentation
- Application of a semantic segmentation convolutional neural network for accurate automatic detection and mapping of solar photovoltaic arrays in aerial imagery
- Video Summarization using Deep Semantic Features
- PrivyNet: A Flexible Framework for Privacy-Preserving Deep Neural Network Training
- Reveal of Domain Effect: How Visual Restoration Contributes to Object Detection in Aquatic Scenes
- Convolutional Neural Networks at Constrained Time Cost
- ZM-Net: Real-time Zero-shot Image Manipulation Network
- Incorporating Global Visual Features into Attention-Based Neural Machine Translation
- Towards End-to-end Text Spotting with Convolutional Recurrent Neural Networks
- Log-DenseNet: How to Sparsify a DenseNet
- A Conditional Generative Model for Predicting Material Microstructures from Processing Methods
- Multi-step Reasoning via Recurrent Dual Attention for Visual Dialog
- RankIQA: Learning from Rankings for No-reference Image Quality Assessment
- Learning to Prune Filters in Convolutional Neural Networks
- Viewpoints and Keypoints
- Joint Object and Part Segmentation using Deep Learned Potentials
- Measurements of Three-Level Hierarchical Structure in the Outliers in the Spectrum of Deepnet Hessians
- Enhance the Motion Cues for Face Anti-Spoofing using CNN-LSTM Architecture
- Unsupervised Representation Learning by Sorting Sequences
- Scribbler: Controlling Deep Image Synthesis with Sketch and Color
- ShrinkTeaNet: Million-scale Lightweight Face Recognition via Shrinking Teacher-Student Networks
- Learning Accurate Low-Bit Deep Neural Networks with Stochastic Quantization
- Deep Feature Consistent Deep Image Transformations: Downscaling, Decolorization and HDR Tone Mapping
- Diversified Texture Synthesis with Feed-forward Networks
- Principled Training of Neural Networks with Direct Feedback Alignment
- One Size Does Not Fit All: Quantifying and Exposing the Accuracy-Latency Trade-off in Machine Learning Cloud Service APIs via Tolerance Tiers
- Deep Neural Networks for Marine Debris Detection in Sonar Images
- Learning Affinity via Spatial Propagation Networks
- Robustness of Neural Networks against Storage Media Errors
- Sliding Line Point Regression for Shape Robust Scene Text Detection
- Simple vs complex temporal recurrences for video saliency prediction
- Modeling Spatial-Temporal Clues in a Hybrid Deep Learning Framework for Video Classification
- Learning Semantic Concepts and Order for Image and Sentence Matching
- VQS: Linking Segmentations to Questions and Answers for Supervised Attention in VQA and Question-Focused Semantic Segmentation
- On Demand Solid Texture Synthesis Using Deep 3D Networks
- Beyond Frontal Faces: Improving Person Recognition Using Multiple Cues
- Automated Cardiothoracic Ratio Calculation and Cardiomegaly Detection using Deep Learning Approach
- Direction Concentration Learning: Enhancing Congruency in Machine Learning
- When Unsupervised Domain Adaptation Meets Tensor Representations
- A Comparative Study of CNN, BoVW and LBP for Classification of Histopathological Images
- Pixie: A System for Recommending 3+ Billion Items to 200+ Million Users in Real-Time
- Temporal Learning and Sequence Modeling for a Job Recommender System
- V2CNet: A Deep Learning Framework to Translate Videos to Commands for Robotic Manipulation
- Attention Based Glaucoma Detection: A Large-scale Database and CNN Model
- Feeding Hand-Crafted Features for Enhancing the Performance of Convolutional Neural Networks
- Black-box Adversarial Attacks with Bayesian Optimization
- Deep Transfer Learning for Static Malware Classification
- Significance-aware Information Bottleneck for Domain Adaptive Semantic Segmentation
- Crafting GBD-Net for Object Detection
- Spatio-Temporal Attention Models for Grounded Video Captioning
- Transfer Learning for Video Recognition with Scarce Training Data for Deep Convolutional Neural Network
- Adaptive NMS: Refining Pedestrian Detection in a Crowd
- SMC Faster R-CNN: Toward a scene-specialized multi-object detector
- FCN-rLSTM: Deep Spatio-Temporal Neural Networks for Vehicle Counting in City Cameras
- What's Mine is Yours: Pretrained CNNs for Limited Training Sonar ATR
- Interpretable Classification from Skin Cancer Histology Slides Using Deep Learning: A Retrospective Multicenter Study
- Fast LIDAR-based Road Detection Using Fully Convolutional Neural Networks
- Histopathologic Image Processing: A Review
- Soft Proposal Networks for Weakly Supervised Object Localization
- Solving Optimization Problems through Fully Convolutional Networks: an Application to the Travelling Salesman Problem
- On the Validity of Bayesian Neural Networks for Uncertainty Estimation
- Dense Scale Network for Crowd Counting
- Robust and High Performance Face Detector
- Residual Features and Unified Prediction Network for Single Stage Detection
- Relevance Prediction from Eye-movements Using Semi-interpretable Convolutional Neural Networks
- GLMNet: Graph Learning-Matching Networks for Feature Matching
- Deep Long Audio Inpainting
- Self-Binarizing Networks
- Peak-Piloted Deep Network for Facial Expression Recognition
- Non-Gaussianity of Stochastic Gradient Noise
- Pattern-Affinitive Propagation across Depth, Surface Normal and Semantic Segmentation
- RED: Reinforced Encoder-Decoder Networks for Action Anticipation
- Learning Uncertain Convolutional Features for Accurate Saliency Detection
- Attention-Based Multimodal Fusion for Video Description
- Dynamic Multi-Task Learning for Face Recognition with Facial Expression
- Factors of Transferability for a Generic ConvNet Representation
- Hierarchical Temporal Convolutional Networks for Dynamic Recommender Systems
- Convolutional Neural Networks for Histopathology Image Classification: Training vs. Using Pre-Trained Networks
- Natural Language Guided Visual Relationship Detection
- Transitive Invariance for Self-supervised Visual Representation Learning
- Hiding Faces in Plain Sight: Disrupting AI Face Synthesis with Adversarial Perturbations
- Residual and Plain Convolutional Neural Networks for 3D Brain MRI Classification
- Towards Compact and Robust Deep Neural Networks
- MarioNETte: Few-shot Face Reenactment Preserving Identity of Unseen Targets
- Task Augmentation by Rotating for Meta-Learning
- W-Net: Reinforced U-Net for Density Map Estimation
- Improving Fully Convolution Network for Semantic Segmentation
- A multi-branch convolutional neural network for detecting double JPEG compression
- An End-to-End Network for Panoptic Segmentation
- Knowledge Concentration: Learning 100K Object Classifiers in a Single CNN
- Task-driven Visual Saliency and Attention-based Visual Question Answering
- Deep CNN-based Multi-task Learning for Open-Set Recognition
- Improving RetinaNet for CT Lesion Detection with Dense Masks from Weak RECIST Labels
- Learning Generalizable and Identity-Discriminative Representations for Face Anti-Spoofing
- Survey of Visual Question Answering: Datasets and Techniques
- Beyond Gaussian Pyramid: Multi-skip Feature Stacking for Action Recognition
- High-Resolution Multispectral Dataset for Semantic Segmentation
- TF-Replicator: Distributed Machine Learning for Researchers
- SIXray : A Large-scale Security Inspection X-ray Benchmark for Prohibited Item Discovery in Overlapping Images
- Towards Interpretable and Robust Hand Detection via Pixel-wise Prediction
- Rules of the Road: Predicting Driving Behavior with a Convolutional Model of Semantic Interactions
- Predicting Aircraft Trajectories: A Deep Generative Convolutional Recurrent Neural Networks Approach
- Bandwidth Extension on Raw Audio via Generative Adversarial Networks
- Hypercolumns for Object Segmentation and Fine-grained Localization
- Fashioning with Networks: Neural Style Transfer to Design Clothes
- A Deep Spatial Contextual Long-term Recurrent Convolutional Network for Saliency Detection
- X-CHANGR: Changing Memristive Crossbar Mapping for Mitigating Line-Resistance Induced Accuracy Degradation in Deep Neural Networks
- Advances in Joint CTC-Attention based End-to-End Speech Recognition with a Deep CNN Encoder and RNN-LM
- Lattice Long Short-Term Memory for Human Action Recognition
- Synthesizing Filamentary Structured Images with GANs
- Learnable Bernoulli Dropout for Bayesian Deep Learning
- An Accurate and Real-time Self-blast Glass Insulator Location Method Based On Faster R-CNN and U-net with Aerial Images
- A new take on measuring relative nutritional density: The feasibility of using a deep neural network to assess commercially-prepared pureed food concentrations
- SkyNet: A Champion Model for DAC-SDC on Low Power Object Detection
- Multi-Scale Attention with Dense Encoder for Handwritten Mathematical Expression Recognition
- Knowledge Adaptation for Efficient Semantic Segmentation
- Material Recognition in the Wild with the Materials in Context Database
- Meta Networks for Neural Style Transfer
- Progressive DARTS: Bridging the Optimization Gap for NAS in the Wild
- Boosting Convolutional Features for Robust Object Proposals
- DeepFlash: Turning a Flash Selfie into a Studio Portrait
- Learning Chained Deep Features and Classifiers for Cascade in Object Detection
- A Scale Invariant Flatness Measure for Deep Network Minima
- From Patches to Pictures (PaQ-2-PiQ): Mapping the Perceptual Space of Picture Quality
- Combating Human Trafficking with Deep Multimodal Models
- There is Limited Correlation between Coverage and Robustness for Deep Neural Networks
- Yet another but more efficient black-box adversarial attack: tiling and evolution strategies
- M2CAI Workflow Challenge: Convolutional Neural Networks with Time Smoothing and Hidden Markov Model for Video Frames Classification
- TorontoCity: Seeing the World with a Million Eyes
- Deep Learning for Object Saliency Detection and Image Segmentation
- Exploit fully automatic low-level segmented PET data for training high-level deep learning algorithms for the corresponding CT data
- Yottixel -- An Image Search Engine for Large Archives of Histopathology Whole Slide Images
- Performance Guaranteed Network Acceleration via High-Order Residual Quantization
- Take it in your stride: Do we need striding in CNNs?
- Effective Use of Dilated Convolutions for Segmenting Small Object Instances in Remote Sensing Imagery
- ShaResNet: reducing residual network parameter number by sharing weights
- Efficient Computation in Adaptive Artificial Spiking Neural Networks
- Attention Clusters: Purely Attention Based Local Feature Integration for Video Classification
- An Adversarial Regularisation for Semi-Supervised Training of Structured Output Neural Networks
- Preserving Patient Privacy while Training a Predictive Model of In-hospital Mortality
- All about Structure: Adapting Structural Information across Domains for Boosting Semantic Segmentation
- CIFAR-10 Image Classification Using Feature Ensembles
- Consistent Optimization for Single-Shot Object Detection
- Neural Person Search Machines
- Learning to Segment Instances in Videos with Spatial Propagation Network
- Efficient Memory Management for GPU-based Deep Learning Systems
- Recognizing American Sign Language Manual Signs from RGB-D Videos
- Learning Social Image Embedding with Deep Multimodal Attention Networks
- Character-level Convolutional Network for Text Classification Applied to Chinese Corpus
- Semantic Redundancies in Image-Classification Datasets: The 10% You Don't Need
- Accelerating Deep Convolutional Networks using low-precision and sparsity
- Joint Iris Segmentation and Localization Using Deep Multi-task Learning Framework
- Are Safer Looking Neighborhoods More Lively? A Multimodal Investigation into Urban Life
- Describing like humans: on diversity in image captioning
- Hybrid Loss for Learning Single-Image-based HDR Reconstruction
- Image Credibility Analysis with Effective Domain Transferred Deep Networks
- Projection Convolutional Neural Networks for 1-bit CNNs via Discrete Back Propagation
- Accurate reconstruction of image stimuli from human fMRI based on the decoding model with capsule network architecture
- Temporal Unet: Sample Level Human Action Recognition using WiFi
- Deep Cuboid Detection: Beyond 2D Bounding Boxes
- Learning Joint Feature Adaptation for Zero-Shot Recognition
- Are State-of-the-art Visual Place Recognition Techniques any Good for Aerial Robotics?
- Multimodal Attribute Extraction
- Active Learning for Visual Question Answering: An Empirical Study
- Zero-Shot Learning with Generative Latent Prototype Model
- Care about you: towards large-scale human-centric visual relationship detection
- Few-Shot Image Recognition by Predicting Parameters from Activations
- II-FCN for skin lesion analysis towards melanoma detection
- Learning Category Correlations for Multi-label Image Recognition with Graph Networks
- Representation-Aggregation Networks for Segmentation of Multi-Gigapixel Histology Images
- Image Aesthetics Assessment Using Composite Features from off-the-Shelf Deep Models
- Chameleon: Adaptive Code Optimization for Expedited Deep Neural Network Compilation
- HashGAN:Attention-aware Deep Adversarial Hashing for Cross Modal Retrieval
- Robust Sparse Regularization: Simultaneously Optimizing Neural Network Robustness and Compactness
- Cats and Captions vs. Creators and the Clock: Comparing Multimodal Content to Context in Predicting Relative Popularity
- Unsupervised Feature Learning with K-means and An Ensemble of Deep Convolutional Neural Networks for Medical Image Classification
- Utilizing the Instability in Weakly Supervised Object Detection
- Half-CNN: A General Framework for Whole-Image Regression
- WeText: Scene Text Detection under Weak Supervision
- Hierarchical LSTMs with Adaptive Attention for Visual Captioning
- Fréchet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms
- ChamNet: Towards Efficient Network Design through Platform-Aware Model Adaptation
- Deep Projective 3D Semantic Segmentation
- Machine learning and AI research for Patient Benefit: 20 Critical Questions on Transparency, Replicability, Ethics and Effectiveness
- Person-in-WiFi: Fine-grained Person Perception using WiFi
- Extend the shallow part of Single Shot MultiBox Detector via Convolutional Neural Network
- DARVIZ: Deep Abstract Representation, Visualization, and Verification of Deep Learning Models
- Audiovisual Transformer Architectures for Large-Scale Classification and Synchronization of Weakly Labeled Audio Events
- An interpretable classifier for high-resolution breast cancer screening images utilizing weakly supervised localization
- Multi-class Classification without Multi-class Labels
- Crafting a multi-task CNN for viewpoint estimation
- Dense Optical Flow based Change Detection Network Robust to Difference of Camera Viewpoints
- Beyond the Pixel-Wise Loss for Topology-Aware Delineation
- Full-Resolution Residual Networks for Semantic Segmentation in Street Scenes
- Learning Multi-Scale Deep Features for High-Resolution Satellite Image Classification
- Privacy-Preserving Deep Inference for Rich User Data on The Cloud
- Knowledge Projection for Deep Neural Networks
- Poison as a Cure: Detecting & Neutralizing Variable-Sized Backdoor Attacks in Deep Neural Networks
- Comprehension-guided referring expressions
- HAMBox: Delving into Online High-quality Anchors Mining for Detecting Outer Faces
- From Deterministic to Generative: Multi-Modal Stochastic RNNs for Video Captioning
- Conservative Wasserstein Training for Pose Estimation
- Are You Talking to Me? Reasoned Visual Dialog Generation through Adversarial Learning
- Learning Pixel-Distribution Prior with Wider Convolution for Image Denoising
- An inner-loop free solution to inverse problems using deep neural networks
- A Generative Adversarial Network for AI-Aided Chair Design
- Discriminatively Boosted Image Clustering with Fully Convolutional Auto-Encoders
- Neural Machine Translation with Latent Semantic of Image and Text
- Bit-Flip Attack: Crushing Neural Network with Progressive Bit Search
- When Semi-Supervised Learning Meets Transfer Learning: Training Strategies, Models and Datasets
- Programmable Neural Network Trojan for Pre-Trained Feature Extractor
- Training on the test set? An analysis of Spampinato et al. [31]
- Generic Feature Learning for Wireless Capsule Endoscopy Analysis
- Graph-Based Classification of Omnidirectional Images
- Justifying Diagnosis Decisions by Deep Neural Networks
- Metamorphic Detection of Adversarial Examples in Deep Learning Models With Affine Transformations
- Grounding Spatio-Semantic Referring Expressions for Human-Robot Interaction
- Adaptive Precision CNN Accelerator Using Radix-X Parallel Connected Memristor Crossbars
- Large Scale Evolution of Convolutional Neural Networks Using Volunteer Computing
- Deep Cross-Modal Correlation Learning for Audio and Lyrics in Music Retrieval
- Additive Noise Annealing and Approximation Properties of Quantized Neural Networks
- DeepNovoV2: Better de novo peptide sequencing with deep learning
- EraseReLU: A Simple Way to Ease the Training of Deep Convolution Neural Networks
- Art of singular vectors and universal adversarial perturbations
- Visually Aligned Word Embeddings for Improving Zero-shot Learning
- Spatio-Temporal Action Detection with Cascade Proposal and Location Anticipation
- Box-driven Class-wise Region Masking and Filling Rate Guided Loss for Weakly Supervised Semantic Segmentation
- Coverage Testing of Deep Learning Models using Dataset Characterization
- A deep learning based solution for construction equipment detection: from development to deployment
- Comparing Neural and Attractiveness-based Visual Features for Artwork Recommendation
- Multilingual Multi-modal Embeddings for Natural Language Processing
- Crime Mapping from Satellite Imagery via Deep Learning
- Unsupervised Deep Features for Privacy Image Classification
- Automatic Classification of Bright Retinal Lesions via Deep Network Features
- Measuring the Transferability of Adversarial Examples
- Contextual Multi-Scale Region Convolutional 3D Network for Activity Detection
- SRDGAN: learning the noise prior for Super Resolution with Dual Generative Adversarial Networks
- A New Convolutional Network-in-Network Structure and Its Applications in Skin Detection, Semantic Segmentation, and Artifact Reduction
- Improving automated multiple sclerosis lesion segmentation with a cascaded 3D convolutional neural network approach
- Ligand Pose Optimization with Atomic Grid-Based Convolutional Neural Networks
- Application of Convolutional Neural Network for Image Classification on Pascal VOC Challenge 2012 dataset
- Video Fill in the Blank with Merging LSTMs
- Deep Mangoes: from fruit detection to cultivar identification in colour images of mango trees
- Reinforcement Learning for Learning Rate Control
- CenterFace: Joint Face Detection and Alignment Using Face as Point
- Luck Matters: Understanding Training Dynamics of Deep ReLU Networks
- Generating lyrics with variational autoencoder and multi-modal artist embeddings
- Contrastive-center loss for deep neural networks
- Synthetic to Real Adaptation with Generative Correlation Alignment Networks
- Inverse Compositional Spatial Transformer Networks
- Quickly Inserting Pegs into Uncertain Holes using Multi-view Images and Deep Network Trained on Synthetic Data
- Neural Networks with Smooth Adaptive Activation Functions for Regression
- A Novel Convolutional Neural Network for Image Steganalysis with Shared Normalization
- Let's Dance: Learning From Online Dance Videos
- Robust Watermarking of Neural Network with Exponential Weighting
- Counting and Segmenting Sorghum Heads
- Deep Active Learning: Unified and Principled Method for Query and Training
- Projection Based Weight Normalization for Deep Neural Networks
- Riemannian approach to batch normalization
- Recurrent Convolutional Networks for Pulmonary Nodule Detection in CT Imaging
- SS-Auto: A Single-Shot, Automatic Structured Weight Pruning Framework of DNNs with Ultra-High Efficiency
- Extracting 3D Vascular Structures from Microscopy Images using Convolutional Recurrent Networks
- KaoKore: A Pre-modern Japanese Art Facial Expression Dataset
- Centripetal SGD for Pruning Very Deep Convolutional Networks with Complicated Structure
- Convolution Neural Network Architecture Learning for Remote Sensing Scene Classification
- Layer-wise Adaptive Gradient Sparsification for Distributed Deep Learning with Convergence Guarantees
- Using Deep Learning Neural Networks and Candlestick Chart Representation to Predict Stock Market
- Learning to Confuse: Generating Training Time Adversarial Data with Auto-Encoder
- Exploring the Design Space of Deep Convolutional Neural Networks at Large Scale
- Cross Domain Knowledge Transfer for Person Re-identification
- Clickbait Identification using Neural Networks
- Fast Recurrent Fully Convolutional Networks for Direct Perception in Autonomous Driving
- Deep GrabCut for Object Selection
- Convolutional Neural Network-Based Image Representation for Visual Loop Closure Detection
- Recurrent Attentional Reinforcement Learning for Multi-label Image Recognition
- DistillHash: Unsupervised Deep Hashing by Distilling Data Pairs
- Frame-Recurrent Video Inpainting by Robust Optical Flow Inference
- Deep Spatial Regression Model for Image Crowd Counting
- Embedded hyper-parameter tuning by Simulated Annealing
- A Revisit on Deep Hashings for Large-scale Content Based Image Retrieval
- Learnable Embedding Space for Efficient Neural Architecture Compression
- Multi-Path Feedback Recurrent Neural Network for Scene Parsing
- Utilizing Deep Learning Towards Multi-modal Bio-sensing and Vision-based Affective Computing
- Classification of Quantitative Light-Induced Fluorescence Images Using Convolutional Neural Network
- COLTRANE: ConvolutiOnaL TRAjectory NEtwork for Deep Map Inference
- Synthesising Dynamic Textures using Convolutional Neural Networks
- An Analysis of Adversarial Attacks and Defenses on Autonomous Driving Models
- Age Group and Gender Estimation in the Wild with Deep RoR Architecture
- Okutama-Action: An Aerial View Video Dataset for Concurrent Human Action Detection
- Modified Distribution Alignment for Domain Adaptation with Pre-trained Inception ResNet
- Visual Discovery at Pinterest
- OpenEI: An Open Framework for Edge Intelligence
- Active Convolution: Learning the Shape of Convolution for Image Classification
- Trained Rank Pruning for Efficient Deep Neural Networks
- Quantifying Facial Age by Posterior of Age Comparisons
- Mis-classified Vector Guided Softmax Loss for Face Recognition
- Example-Guided Style Consistent Image Synthesis from Semantic Labeling
- Cross-domain Human Parsing via Adversarial Feature and Label Adaptation
- Prediction of Kidney Function from Biopsy Images Using Convolutional Neural Networks
- Machine Vision for Natural Gas Methane Emissions Detection Using an Infrared Camera
- Tell-and-Answer: Towards Explainable Visual Question Answering using Attributes and Captions
- Single-Network Whole-Body Pose Estimation
- What Looks Good with my Sofa: Multimodal Search Engine for Interior Design
- Efficient Segmentation: Learning Downsampling Near Semantic Boundaries
- SREdgeNet: Edge Enhanced Single Image Super Resolution using Dense Edge Detection Network and Feature Merge Network
- Musical Tempo and Key Estimation using Convolutional Neural Networks with Directional Filters
- Verification of Very Low-Resolution Faces Using An Identity-Preserving Deep Face Super-Resolution Network
- Multisource and Multitemporal Data Fusion in Remote Sensing
- Fine-Grained Neural Architecture Search
- Wide and deep volumetric residual networks for volumetric image classification
- Improving Device-Edge Cooperative Inference of Deep Learning via 2-Step Pruning
- Deep Residual Networks and Weight Initialization
- Detecting Deepfake-Forged Contents with Separable Convolutional Neural Network and Image Segmentation
- HyperNetworks with statistical filtering for defending adversarial examples
- CN-CELEB: a challenging Chinese speaker recognition dataset
- A Survey on Biomedical Image Captioning
- Continual Learning in Neural Networks
- A Coarse-to-Fine Adaptive Network for Appearance-Based Gaze Estimation
- Spatio-Temporal Fusion Networks for Action Recognition
- The Application of Two-level Attention Models in Deep Convolutional Neural Network for Fine-grained Image Classification
- A Comparison of deep learning methods for environmental sound
- Deep Heterogeneous Feature Fusion for Template-Based Face Recognition
- Modularized Morphing of Neural Networks
- Machine Learning for Dental Image Analysis
- Very Deep Convolutional Neural Networks for Robust Speech Recognition
- Learning Common and Specific Features for RGB-D Semantic Segmentation with Deconvolutional Networks
- Learning deep representation from coarse to fine for face alignment
- Dynamic Optimization of Neural Network Structures Using Probabilistic Modeling
- Structured Inhomogeneous Density Map Learning for Crowd Counting
- BT-Nets: Simplifying Deep Neural Networks via Block Term Decomposition
- Deep Learning for identifying radiogenomic associations in breast cancer
- Deep Matching Autoencoders
- Classification of Radio Signals and HF Transmission Modes with Deep Learning
- A Scalable Learned Index Scheme in Storage Systems
- TAFE-Net: Task-Aware Feature Embeddings for Low Shot Learning
- Dense Intrinsic Appearance Flow for Human Pose Transfer
- Data Augmentation with Manifold Exploring Geometric Transformations for Increased Performance and Robustness
- Distance-Based Learning from Errors for Confidence Calibration
- Deep Sub-Ensembles for Fast Uncertainty Estimation in Image Classification
- Unsupervised Medical Image Segmentation with Adversarial Networks: From Edge Diagrams to Segmentation Maps
- ModelHub: Towards Unified Data and Lifecycle Management for Deep Learning
- C3AE: Exploring the Limits of Compact Model for Age Estimation
- MHTN: Modal-adversarial Hybrid Transfer Network for Cross-modal Retrieval
- Training Sparse Neural Networks
- BLK-REW: A Unified Block-based DNN Pruning Framework using Reweighted Regularization Method
- Deep Image Category Discovery using a Transferred Similarity Function
- Transferability of Adversarial Examples to Attack Cloud-based Image Classifier Service
- Deep Spiking Neural Network with Spike Count based Learning Rule
- Proximal Alternating Direction Network: A Globally Converged Deep Unrolling Framework
- Actions and Attributes from Wholes and Parts
- Integrating both Visual and Audio Cues for Enhanced Video Caption
- Unsupervised Pose Flow Learning for Pose Guided Synthesis
- EdgeCNN: Convolutional Neural Network Classification Model with small inputs for Edge Computing
- Single Image Super Resolution - When Model Adaptation Matters
- Class Rectification Hard Mining for Imbalanced Deep Learning
- Context-Aware Image Matting for Simultaneous Foreground and Alpha Estimation
- Deep Kernel Learning via Random Fourier Features
- StartNet: Online Detection of Action Start in Untrimmed Videos
- Guiding Neuroevolution with Structural Objectives
- Predicting Food Security Outcomes Using Convolutional Neural Networks (CNNs) for Satellite Tasking
- Recent Advances in the Applications of Convolutional Neural Networks to Medical Image Contour Detection
- An End-to-End Approach to Natural Language Object Retrieval via Context-Aware Deep Reinforcement Learning
- Bayesian Learning of Neural Network Architectures
- Deep Learning for Target Classification from SAR Imagery: Data Augmentation and Translation Invariance
- Sequential Dual Deep Learning with Shape and Texture Features for Sketch Recognition
- The Role of Context Selection in Object Detection
- Deep Co-Space: Sample Mining Across Feature Transformation for Semi-Supervised Learning
- Adversarial nets with perceptual losses for text-to-image synthesis
- Low-Precision Batch-Normalized Activations
- Dual-branch residual network for lung nodule segmentation
- Training Deep Networks for Facial Expression Recognition with Crowd-Sourced Label Distribution
- Cystoid macular edema segmentation of Optical Coherence Tomography images using fully convolutional neural networks and fully connected CRFs
- Deep Learning in Wide-field Surveys: Fast Analysis of Strong Lenses in Ground-based Cosmic Experiments
- Is Saki #delicious? The Food Perception Gap on Instagram and Its Relation to Health
- Estimated Depth Map Helps Image Classification
- CGaP: Continuous Growth and Pruning for Efficient Deep Learning
- Exploiting Human Social Cognition for the Detection of Fake and Fraudulent Faces via Memory Networks
- RAN4IQA: Restorative Adversarial Nets for No-Reference Image Quality Assessment
- Uncertainty-Aware Driver Trajectory Prediction at Urban Intersections
- Semantically Consistent Image Completion with Fine-grained Details
- Towards Compact ConvNets via Structure-Sparsity Regularized Filter Pruning
- Multiple Instance Curriculum Learning for Weakly Supervised Object Detection
- Deep Matching and Validation Network -- An End-to-End Solution to Constrained Image Splicing Localization and Detection
- ADMM-NN: An Algorithm-Hardware Co-Design Framework of DNNs Using Alternating Direction Method of Multipliers
- Sigma Delta Quantized Networks
- ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering
- Deep Restricted Boltzmann Networks
- A Review of Meta-Reinforcement Learning for Deep Neural Networks Architecture Search
- Convolutional Residual Memory Networks
- Learning RGB-D Salient Object Detection using background enclosure, depth contrast, and top-down features
- Positively Scale-Invariant Flatness of ReLU Neural Networks
- SESR: Single Image Super Resolution with Recursive Squeeze and Excitation Networks
- Attention Allocation Aid for Visual Search
- Lending Orientation to Neural Networks for Cross-view Geo-localization
- Deep Exemplar-based Video Colorization
- A Bi-Directional Co-Design Approach to Enable Deep Learning on IoT Devices
- High Resolution Millimeter Wave Imaging For Self-Driving Cars
- Image Quality Assessment Guided Deep Neural Networks Training
- PI-REC: Progressive Image Reconstruction Network With Edge and Color Domain
- Fast and Efficient Zero-Learning Image Fusion
- ScanNet: A Fast and Dense Scanning Framework for Metastatic Breast Cancer Detection from Whole-Slide Images
- HyperNOMAD: Hyperparameter optimization of deep neural networks using mesh adaptive direct search
- Attention is all you need for Videos: Self-attention based Video Summarization using Universal Transformers
- Dual Path Networks for Multi-Person Human Pose Estimation
- A Way out of the Odyssey: Analyzing and Combining Recent Insights for LSTMs
- A Two-stream End-to-End Deep Learning Network for Recognizing Atypical Visual Attention in Autism Spectrum Disorder
- Cross-Domain Cascaded Deep Feature Translation
- Visual-Textual Association with Hardest and Semi-Hard Negative Pairs Mining for Person Search
- 3DN: 3D Deformation Network
- Active Learning with TensorBoard Projector
- Anti-Makeup: Learning A Bi-Level Adversarial Network for Makeup-Invariant Face Verification
- Learning a Repression Network for Precise Vehicle Search
- A Dilated Inception Network for Visual Saliency Prediction
- Holistic Interstitial Lung Disease Detection using Deep Convolutional Neural Networks: Multi-label Learning and Unordered Pooling
- Cultural Event Recognition with Visual ConvNets and Temporal Models
- Generative Adversarial Network based on Resnet for Conditional Image Restoration
- Learning Structured Semantic Embeddings for Visual Recognition
- Fingerprint Spoof Generalization
- Collaborative Summarization of Topic-Related Videos
- Deep Learning Models for Digital Pathology
- NESTA: Hamming Weight Compression-Based Neural Proc. Engine
- Vid2Game: Controllable Characters Extracted from Real-World Videos
- AFP-Net: Realtime Anchor-Free Polyp Detection in Colonoscopy
- Unconstrained Scene Text and Video Text Recognition for Arabic Script
- See, Attend and Brake: An Attention-based Saliency Map Prediction Model for End-to-End Driving
- Tensor Contraction Layers for Parsimonious Deep Nets
- The Color of the Cat is Gray: 1 Million Full-Sentences Visual Question Answering (FSVQA)
- Gradient Regularization for Quantization Robustness
- Wavelet Domain Style Transfer for an Effective Perception-distortion Tradeoff in Single Image Super-Resolution
- Differential Generative Adversarial Networks: Synthesizing Non-linear Facial Variations with Limited Number of Training Data
- cGAN-based Manga Colorization Using a Single Training Image
- Free-form Video Inpainting with 3D Gated Convolution and Temporal PatchGAN
- A Comprehensive Study of Alzheimer's Disease Classification Using Convolutional Neural Networks
- Detailed Human Shape Estimation from a Single Image by Hierarchical Mesh Deformation
- Learning language through pictures
- IoU-uniform R-CNN: Breaking Through the Limitations of RPN
- Geometric robustness of deep networks: analysis and improvement
- Alzheimer's Disease Brain MRI Classification: Challenges and Insights
- Neural Inverse Knitting: From Images to Manufacturing Instructions
- CUP: Cluster Pruning for Compressing Deep Neural Networks
- Segmental Convolutional Neural Networks for Detection of Cardiac Abnormality With Noisy Heart Sound Recordings
- Per-Tensor Fixed-Point Quantization of the Back-Propagation Algorithm
- Self-explanatory Deep Salient Object Detection
- A SOT-MRAM-based Processing-In-Memory Engine for Highly Compressed DNN Implementation
- Adaptive Feeding: Achieving Fast and Accurate Detections by Adaptively Combining Object Detectors
- A 4D Light-Field Dataset and CNN Architectures for Material Recognition
- Fully-adaptive Feature Sharing in Multi-Task Networks with Applications in Person Attribute Classification
- Automatic Information Extraction from Piping and Instrumentation Diagrams
- Deep Learning Based Automatic Video Annotation Tool for Self-Driving Car
- Relating Input Concepts to Convolutional Neural Network Decisions
- Zero-shot Learning via Shared-Reconstruction-Graph Pursuit
- DwNet: Dense warp-based network for pose-guided human video generation
- What Is the Best Practice for CNNs Applied to Visual Instance Retrieval?
- GASL: Guided Attention for Sparsity Learning in Deep Neural Networks
- Fully Automatic Segmentation of Lumbar Vertebrae from CT Images using Cascaded 3D Fully Convolutional Networks
- Enabling Privacy-Preserving, Compute- and Data-Intensive Computing using Heterogeneous Trusted Execution Environment
- 3D Reconstruction in Canonical Co-ordinate Space from Arbitrarily Oriented 2D Images
- Learning Pairwise Relationship for Multi-object Detection in Crowded Scenes
- MUSEFood: Multi-sensor-based Food Volume Estimation on Smartphones
- Neuralogram: A Deep Neural Network Based Representation for Audio Signals
- Toward Streaming Synapse Detection with Compositional ConvNets
- DeepFeat: A Bottom Up and Top Down Saliency Model Based on Deep Features of Convolutional Neural Nets
- Leveraging Medical Visual Question Answering with Supporting Facts
- ModelHub.AI: Dissemination Platform for Deep Learning Models
- A Distributed Synchronous SGD Algorithm with Global Top- Sparsification for Low Bandwidth Networks
- Filter Grafting for Deep Neural Networks
- Robust or Private? Adversarial Training Makes Models More Vulnerable to Privacy Attacks
- Adversarial Perturbations Prevail in the Y-Channel of the YCbCr Color Space
- Neural Rerendering in the Wild
- ViWi Vision-Aided mmWave Beam Tracking: Dataset, Task, and Baseline Solutions
- Temporal Segmentation of Surgical Sub-tasks through Deep Learning with Multiple Data Sources
- Seismic data denoising and deblending using deep learning
- An Empirical Exploration of Skip Connections for Sequential Tagging
- Deep TEN: Texture Encoding Network
- Camera Lens Super-Resolution
- Blockchain-based Bidirectional Updates on Fine-grained Medical Data
- A New Loss Function for CNN Classifier Based on Pre-defined Evenly-Distributed Class Centroids
- Survey of Recent Advances in Visual Question Answering
- Pre-Trained Convolutional Neural Network Features for Facial Expression Recognition
- Improving Facial Emotion Recognition Systems Using Gradient and Laplacian Images
- Attacking Automatic Video Analysis Algorithms: A Case Study of Google Cloud Video Intelligence API
- Deep Structured-Output Regression Learning for Computational Color Constancy
- Tree-structured Kronecker Convolutional Network for Semantic Segmentation
- Deep Learning applied to Road Traffic Speed forecasting
- Pixel-aware Deep Function-mixture Network for Spectral Super-Resolution
- An Application-Specific VLIW Processor with Vector Instruction Set for CNN Acceleration
- Reliable and Efficient Image Cropping: A Grid Anchor based Approach
- SWNet: Small-World Neural Networks and Rapid Convergence
- Deep Learning Based Computed Tomography Whys and Wherefores
- FOTS: Fast Oriented Text Spotting with a Unified Network
- Accelerated Training for Massive Classification via Dynamic Class Selection
- PaintBot: A Reinforcement Learning Approach for Natural Media Painting
- Optimized Broadcast for Deep Learning Workloads on Dense-GPU InfiniBand Clusters: MPI or NCCL?
- Deep learning with spatiotemporal consistency for nerve segmentation in ultrasound images
- Texture segmentation with Fully Convolutional Networks
- Neural Networks Regularization Through Class-wise Invariant Representation Learning
- DualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual Dialogue
- Deep Learning the Indus Script
- Translation Insensitive CNNs
- Multi-Task Driven Feature Models for Thermal Infrared Tracking
- Bidirectional Scene Text Recognition with a Single Decoder
- A concatenating framework of shortcut convolutional neural networks
- E2-Capsule Neural Networks for Facial Expression Recognition Using AU-Aware Attention
- Monocular Depth Estimation with Hierarchical Fusion of Dilated CNNs and Soft-Weighted-Sum Inference
- Automatic quality assessment for 2D fetal sonographic standard plane based on multi-task learning
- Autoencoder-Based Incremental Class Learning without Retraining on Old Data
- Vid2speech: Speech Reconstruction from Silent Video
- A Deep Neuro-Fuzzy Network for Image Classification
- Towards Characterizing and Limiting Information Exposure in DNN Layers
- ALFA: Agglomerative Late Fusion Algorithm for Object Detection
- A Systematic Mapping Study on Testing of Machine Learning Programs
- An Efficient Hardware Accelerator for Structured Sparse Convolutional Neural Networks on FPGAs
- TrackNet: A Deep Learning Network for Tracking High-speed and Tiny Objects in Sports Applications
- Contextual Visual Similarity
- NETNet: Neighbor Erasing and Transferring Network for Better Single Shot Object Detection
- Photographic home styles in Congress: a computer vision approach
- Training Bit Fully Convolutional Network for Fast Semantic Segmentation
- SIP-SegNet: A Deep Convolutional Encoder-Decoder Network for Joint Semantic Segmentation and Extraction of Sclera, Iris and Pupil based on Periocular Region Suppression
- Face Parsing via Recurrent Propagation
- Audio-video Emotion Recognition in the Wild using Deep Hybrid Networks
- I Am Going MAD: Maximum Discrepancy Competition for Comparing Classifiers Adaptively
- Disentangling Adaptive Gradient Methods from Learning Rates
- MVP: Unified Motion and Visual Self-Supervised Learning for Large-Scale Robotic Navigation
- Privacy-preserving Learning via Deep Net Pruning
- Multimodal Memory Modelling for Video Captioning
- Kiki Kills: Identifying Dangerous Challenge Videos from Social Media
- Structured Pruning for Efficient ConvNets via Incremental Regularization
- Wasserstein Style Transfer
- MAVOT: Memory-Augmented Video Object Tracking
- Incremental Learning Using a Grow-and-Prune Paradigm with Efficient Neural Networks
- Weightless Neural Network with Transfer Learning to Detect Distress in Asphalt
- Transformed Regularization for Learning Sparse Deep Neural Networks
- Convolutional Neural Networks on non-uniform geometrical signals using Euclidean spectral transformation
- A Curriculum Domain Adaptation Approach to the Semantic Segmentation of Urban Scenes
- Towards Testing of Deep Learning Systems with Training Set Reduction
- Object-centric Sampling for Fine-grained Image Classification
- Test Selection for Deep Learning Systems
- Gated Context Model with Embedded Priors for Deep Image Compression
- Finding ReMO (Related Memory Object): A Simple Neural Architecture for Text based Reasoning
- Improve Object Detection by Data Enhancement based on Generative Adversarial Nets
- Deep CNNs along the Time Axis with Intermap Pooling for Robustness to Spectral Variations
- A Theoretically Sound Upper Bound on the Triplet Loss for Improving the Efficiency of Deep Distance Metric Learning
- Affordance Learning In Direct Perception for Autonomous Driving
- On the Mathematical Understanding of ResNet with Feynman Path Integral
- Selective Deep Convolutional Features for Image Retrieval
- Learning Whole-Image Descriptors for Real-time Loop Detection andKidnap Recovery under Large Viewpoint Difference
- An Out-of-the-box Full-network Embedding for Convolutional Neural Networks
- Deep Fusion: An Attention Guided Factorized Bilinear Pooling for Audio-video Emotion Recognition
- Softmax-based Classification is k-means Clustering: Formal Proof, Consequences for Adversarial Attacks, and Improvement through Centroid Based Tailoring
- Virtual Piano using Computer Vision
- Beyond Fine Tuning: A Modular Approach to Learning on Small Data
- Is Faster R-CNN Doing Well for Pedestrian Detection?
- Automatically Searching for U-Net Image Translator Architecture
- Towards Scalable, Efficient and Accurate Deep Spiking Neural Networks with Backward Residual Connections, Stochastic Softmax and Hybridization
- DCNN-GAN: Reconstructing Realistic Image from fMRI
- An ELU Network with Total Variation for Image Denoising
- Co-PACRR: A Context-Aware Neural IR Model for Ad-hoc Retrieval
- GANosaic: Mosaic Creation with Generative Texture Manifolds
- Constrained deep neural network architecture search for IoT devices accounting hardware calibration
- An Adversarial Perturbation Oriented Domain Adaptation Approach for Semantic Segmentation
- Local Label Propagation for Large-Scale Semi-Supervised Learning
- DAmageNet: A Universal Adversarial Dataset
- Autoencoder Based Residual Deep Networks for Robust Regression Prediction and Spatiotemporal Estimation
- A Novel Automation-Assisted Cervical Cancer Reading Method Based on Convolutional Neural Network
- Better Text Understanding Through Image-To-Text Transfer
- Very Deep Convolutional Neural Networks for Raw Waveforms
- Social Scene Understanding: End-to-End Multi-Person Action Localization and Collective Activity Recognition
- SiCloPe: Silhouette-Based Clothed People
- On Classification of Distorted Images with Deep Convolutional Neural Networks
- Discriminative convolutional Fisher vector network for action recognition
- Surveillance Video Parsing with Single Frame Supervision
- DAVE: A Unified Framework for Fast Vehicle Detection and Annotation
- Vision-based Robot Manipulation Learning via Human Demonstrations
- Photo Filter Recommendation by Category-Aware Aesthetic Learning
- ZipNet-GAN: Inferring Fine-grained Mobile Traffic Patterns via a Generative Adversarial Neural Network
- Emotion Recognition for In-the-wild Videos
- Compressibility Loss for Neural Network Weights
- Smart Library: Identifying Books in a Library using Richly Supervised Deep Scene Text Reading
- Better the Devil you Know: An Analysis of Evasion Attacks using Out-of-Distribution Adversarial Examples
- Deep Blind Image Inpainting
- A Simple General Approach to Balance Task Difficulty in Multi-Task Learning
- High-Performance Deep Learning via a Single Building Block
- Exploiting the Tradeoff between Program Accuracy and Soft-error Resiliency Overhead for Machine Learning Workloads
- Detecting Semantic Parts on Partially Occluded Objects
- Semi-Supervised Haptic Material Recognition for Robots using Generative Adversarial Networks
- Context-aware stacked convolutional neural networks for classification of breast carcinomas in whole-slide histopathology images
- Single-Camera Basketball Tracker through Pose and Semantic Feature Fusion
- Genetic Programming and Gradient Descent: A Memetic Approach to Binary Image Classification
- Object-Scene Convolutional Neural Networks for Event Recognition in Images
- Deep Scene Text Detection with Connected Component Proposals
- Towards Proof Synthesis Guided by Neural Machine Translation for Intuitionistic Propositional Logic
- Real-world Mapping of Gaze Fixations Using Instance Segmentation for Road Construction Safety Applications
- Chinese Typeface Transformation with Hierarchical Adversarial Network
- Multi-Stage Pathological Image Classification using Semantic Segmentation
- Be Like Water: Robustness to Extraneous Variables Via Adaptive Feature Normalization
- Teaching DNNs to design fast fashion
- Deep Variational Sufficient Dimensionality Reduction
- Acoustic Scene Classification Using Fusion of Attentive Convolutional Neural Networks for DCASE2019 Challenge
- Deep Stacked Networks with Residual Polishing for Image Inpainting
- DSNet: An Efficient CNN for Road Scene Segmentation
- Revealing structure components of the retina by deep learning networks
- Learning a CNN-based End-to-End Controller for a Formula SAE Racecar
- Context-Aware Semantic Inpainting
- Unconstrained Facial Expression Transfer using Style-based Generator
- A Neural Spiking Approach Compared to Deep Feedforward Networks on Stepwise Pixel Erasement
- Challenges in Time-Stamp Aware Anomaly Detection in Traffic Videos
- Cross-Domain Face Verification: Matching ID Document and Self-Portrait Photographs
- Micro-Expression Spotting: A Benchmark
- Making CNNs for Video Parsing Accessible
- Single Image Super-Resolution Using Multi-Scale Convolutional Neural Network
- A Novel Stochastic Stratified Average Gradient Method: Convergence Rate and Its Complexity
- Multi-Modal Attention-based Fusion Model for Semantic Segmentation of RGB-Depth Images
- Deep residual learning in CT physics: scatter correction for spectral CT
- AI-Skin : Skin Disease Recognition based on Self-learning and Wide Data Collection through a Closed Loop Framework
- How hard can it be? Estimating the difficulty of visual search in an image
- SuperChat: Dialogue Generation by Transfer Learning from Vision to Language using Two-dimensional Word Embedding and Pretrained ImageNet CNN Models
- Scaling Binarized Neural Networks on Reconfigurable Logic
- A dataset and exploration of models for understanding video data through fill-in-the-blank question-answering
- Shadow Transfer: Single Image Relighting For Urban Road Scenes
- DeepOBS: A Deep Learning Optimizer Benchmark Suite
- MMKG: Multi-Modal Knowledge Graphs
- Simultaneous Region Localization and Hash Coding for Fine-grained Image Retrieval
- A3GAN: An Attribute-aware Attentive Generative Adversarial Network for Face Aging
- Face Detection, Bounding Box Aggregation and Pose Estimation for Robust Facial Landmark Localisation in the Wild
- Deep Learning Assessment of Tumor Proliferation in Breast Cancer Histological Images
- Dynamically Visual Disambiguation of Keyword-based Image Search
- DENet: A Universal Network for Counting Crowd with Varying Densities and Scales
- Providing theoretical learning guarantees to Deep Learning Networks
- ArbiText: Arbitrary-Oriented Text Detection in Unconstrained Scene
- Deep Learning for Metagenomic Data: using 2D Embeddings and Convolutional Neural Networks
- Stochastic Downsampling for Cost-Adjustable Inference and Improved Regularization in Convolutional Networks
- InverseNet: Solving Inverse Problems with Splitting Networks
- A Fast and Compact Saliency Score Regression Network Based on Fully Convolutional Network
- Real-time Memory Efficient Large-pose Face Alignment via Deep Evolutionary Network
- U-Net Based Multi-instance Video Object Segmentation
- Learning Perspective Undistortion of Portraits
- Computation Error Analysis of Block Floating Point Arithmetic Oriented Convolution Neural Network Accelerator Design
- Chain-NN: An Energy-Efficient 1D Chain Architecture for Accelerating Deep Convolutional Neural Networks
- NeuNetS: An Automated Synthesis Engine for Neural Network Design
- Object classification in images of Neoclassical furniture using Deep Learning
- Self-Supervised Visual Place Recognition Learning in Mobile Robots
- A Stochastic Extra-Step Quasi-Newton Method for Nonsmooth Nonconvex Optimization
- CNN-based Automatic Detection of Bone Conditions via Diagnostic CT Images for Osteoporosis Screening
- Multi-stage Object Detection with Group Recursive Learning
- Deep execution monitor for robot assistive tasks
- Recognizing Material Properties from Images
- Error Correction for Dense Semantic Image Labeling
- Multi-Prototype Networks for Unconstrained Set-based Face Recognition
- ResNet Can Be Pruned 60x: Introducing Network Purification and Unused Path Removal (P-RM) after Weight Pruning
- Learning Correspondence from the Cycle-Consistency of Time
- Towards Structured Analysis of Broadcast Badminton Videos
- Breast mass segmentation based on ultrasonic entropy maps and attention gated U-Net
- Semantic Relationships Guided Representation Learning for Facial Action Unit Recognition
- Multi-Level Recurrent Residual Networks for Action Recognition
- Building Data-driven Models with Microstructural Images: Generalization and Interpretability
- Classifying the classifier: dissecting the weight space of neural networks
- COP: Customized Deep Model Compression via Regularized Correlation-Based Filter-Level Pruning
- Indexical Cities: Articulating Personal Models of Urban Preference with Geotagged Data
- On the Diversity of Realistic Image Synthesis
- Deep Connectomics Networks: Neural Network Architectures Inspired by Neuronal Networks
- Deep Fusion of Local and Non-Local Features for Precision Landslide Recognition
- On architectural choices in deep learning: From network structure to gradient convergence and parameter estimation
- Co-Regularized Deep Representations for Video Summarization
- Super-resolution Using Constrained Deep Texture Synthesis
- Action-Driven Object Detection with Top-Down Visual Attentions
- Deep Convolutional Neural Network for 6-DOF Image Localization
- DuBox: No-Prior Box Objection Detection via Residual Dual Scale Detectors
- STN-Homography: estimate homography parameters directly
- Deriving Emotions and Sentiments from Visual Content: A Disaster Analysis Use Case
- Large-scale Isolated Gesture Recognition Using Convolutional Neural Networks
- AI Oriented Large-Scale Video Management for Smart City: Technologies, Standards and Beyond
- VIPLFaceNet: An Open Source Deep Face Recognition SDK
- PMC-GANs: Generating Multi-Scale High-Quality Pedestrian with Multimodal Cascaded GANs
- Pattern Generation Strategies for Improving Recognition of Handwritten Mathematical Expressions
- Detecting retail products in situ using CNN without human effort labeling
- Hierarchical Multi-scale Attention Networks for Action Recognition
- Commonly Uncommon: Semantic Sparsity in Situation Recognition
- A Pre-defined Sparse Kernel Based Convolution for Deep CNNs
- Digital Passport: A Novel Technological Strategy for Intellectual Property Protection of Convolutional Neural Networks
- Deep Learning in Robotics: A Review of Recent Research
- Generalizing Deep Models for Overhead Image Segmentation Through Getis-Ord Gi* Pooling
- Improving Direct Physical Properties Prediction of Heterogeneous Materials from Imaging Data via Convolutional Neural Network and a Morphology-Aware Generative Model
- Human perception in computer vision
- Measuring and Understanding Sensory Representations within Deep Networks Using a Numerical Optimization Framework
- Searching Action Proposals via Spatial Actionness Estimation and Temporal Path Inference and Tracking
- Distantly Supervised Road Segmentation
- Auxiliary Multimodal LSTM for Audio-visual Speech Recognition and Lipreading
- Semantic-aware Image Deblurring
- Detection and Attention: Diagnosing Pulmonary Lung Cancer from CT by Imitating Physicians
- Kernelized Wasserstein Natural Gradient
- Deeply-Supervised Recurrent Convolutional Neural Network for Saliency Detection
- A Simple Saliency Method That Passes the Sanity Checks
- A Flow Model of Neural Networks
- Real-Time Steganalysis for Stream Media Based on Multi-channel Convolutional Sliding Windows
- Mixed-Precision Quantized Neural Network with Progressively Decreasing Bitwidth For Image Classification and Object Detection
- Image Captioning with Object Detection and Localization
- DGCNN: Disordered Graph Convolutional Neural Network Based on the Gaussian Mixture Model
- Data Sanity Check for Deep Learning Systems via Learnt Assertions
- Neural Abstract Style Transfer for Chinese Traditional Painting
- Full-Network Embedding in a Multimodal Embedding Pipeline
- Balance Between Efficient and Effective Learning: Dense2Sparse Reward Shaping for Robot Manipulation with Environment Uncertainty
- Perceptually Motivated Method for Image Inpainting Comparison
- Large-Scale 3D Scene Classification With Multi-View Volumetric CNN
- Accurate Automatic Segmentation of Amygdala Subnuclei and Modeling of Uncertainty via Bayesian Fully Convolutional Neural Network
- BusyHands: A Hand-Tool Interaction Database for Assembly Tasks Semantic Segmentation
- On the effect of age perception biases for real age regression
- Mechanisms of Artistic Creativity in Deep Learning Neural Networks
- Sionnx: Automatic Unit Test Generator for ONNX Conformance
- DRAMNet: Authentication based on Physical Unique Features of DRAM Using Deep Convolutional Neural Networks
- Phoenix: A Low-Precision Floating-Point Quantization Oriented Architecture for Convolutional Neural Networks
- Structured Prediction using cGANs with Fusion Discriminator
- TextCohesion: Detecting Text for Arbitrary Shapes
- A Novel Monocular Disparity Estimation Network with Domain Transformation and Ambiguity Learning
- Attend in groups: a weakly-supervised deep learning framework for learning from web data
- Deep Learning in the Automotive Industry: Recent Advances and Application Examples
- Deep learning in bioinformatics: introduction, application, and perspective in big data era
- High dynamic range image forensics using cnn
- Compressing complex convolutional neural network based on an improved deep compression algorithm
- Channel Equilibrium Networks for Learning Deep Representation
- Hybrid Composition with IdleBlock: More Efficient Networks for Image Recognition
- Eliminating artefacts in Polarimetric Images using Deep Learning
- The Tensor Brain: Semantic Decoding for Perception and Memory
- Lung CT Imaging Sign Classification through Deep Learning on Small Data
- NAIS: Neural Architecture and Implementation Search and its Applications in Autonomous Driving
- Learning to Synthesize Fashion Textures
- Efficient Dense Labeling of Human Activity Sequences from Wearables using Fully Convolutional Networks
- EDIT: Exemplar-Domain Aware Image-to-Image Translation
- Can a CNN Recognize Catalan Diet?
- Controllable Descendant Face Synthesis
- Deep Learning and Control Algorithms of Direct Perception for Autonomous Driving
- Bridging the Gap Between Neural Networks and Neuromorphic Hardware with A Neural Network Compiler
- Image Super-Resolution via RL-CSC: When Residual Learning Meets Convolutional Sparse Coding
- When Fashion Meets Big Data: Discriminative Mining of Best Selling Clothing Features
- Structural Pruning in Deep Neural Networks: A Small-World Approach
- MILDNet: A Lightweight Single Scaled Deep Ranking Architecture
- GestARLite: An On-Device Pointing Finger Based Gestural Interface for Smartphones and Video See-Through Head-Mounts
- Pano2CAD: Room Layout From A Single Panorama Image
- Towards End-to-End Audio-Sheet-Music Retrieval
- Beyond Forward Shortcuts: Fully Convolutional Master-Slave Networks (MSNets) with Backward Skip Connections for Semantic Segmentation
- Unsupervised Enhancement of Real-World Depth Images Using Tri-Cycle GAN
- SparCE: Sparsity aware General Purpose Core Extensions to Accelerate Deep Neural Networks
- Deep Features for Tissue-Fold Detection in Histopathology Images
- Distortion Agnostic Deep Watermarking
- Bidirectional Beam Search: Forward-Backward Inference in Neural Sequence Models for Fill-in-the-Blank Image Captioning
- System Demo for Transfer Learning across Vision and Text using Domain Specific CNN Accelerator for On-Device NLP Applications
- Visual Stability Prediction and Its Application to Manipulation
- ORIGAMI: A Heterogeneous Split Architecture for In-Memory Acceleration of Learning
- Group-wise Deep Co-saliency Detection
- A geometry-inspired decision-based attack
- Deep Learning Approach for Very Similar Objects Recognition Application on Chihuahua and Muffin Problem
- Rank of Experts: Detection Network Ensemble
- Learning long-term dependencies for action recognition with a biologically-inspired deep network
- Automatic Mass Detection in Breast Using Deep Convolutional Neural Network and SVM Classifier
- Adaptive Precision Training: Quantify Back Propagation in Neural Networks with Fixed-point Numbers
- Learning a Robust Representation via a Deep Network on Symmetric Positive Definite Manifolds
- Compressed Learning of Deep Neural Networks for OpenCL-Capable Embedded Systems
- Face Alignment In-the-Wild: A Survey
- Binarized Convolutional Neural Networks with Separable Filters for Efficient Hardware Acceleration
- Probabilistic Neural Network with Complex Exponential Activation Functions in Image Recognition using Deep Learning Framework
- Building Graph Representations of Deep Vector Embeddings
- Summarization of ICU Patient Motion from Multimodal Multiview Videos
- Brain Abnormality Detection by Deep Convolutional Neural Network
- Large-Scale YouTube-8M Video Understanding with Deep Neural Networks
- Recovering Homography from Camera Captured Documents using Convolutional Neural Networks
- Classification and Retrieval of Digital Pathology Scans: A New Dataset
- Single Image Action Recognition by Predicting Space-Time Saliency
- Energy-efficient Amortized Inference with Cascaded Deep Classifiers
- Open DNN Box by Power Side-Channel Attack
- Video Action Recognition Via Neural Architecture Searching
- Fast Universal Style Transfer for Artistic and Photorealistic Rendering
- Dissimilarity learning via Siamese network predicts brain imaging data
- Deep Learning from Noisy Image Labels with Quality Embedding
- Phrase-based Image Captioning with Hierarchical LSTM Model
- LDMNet: Low Dimensional Manifold Regularized Neural Networks
- Enhanced Attacks on Defensively Distilled Deep Neural Networks
- Efficient Project Gradient Descent for Ensemble Adversarial Attack
- Label Universal Targeted Attack
- Model Similarity Mitigates Test Set Overuse
- TRk-CNN: Transferable Ranking-CNN for image classification of glaucoma, glaucoma suspect, and normal eyes
- FH-GAN: Face Hallucination and Recognition using Generative Adversarial Network
- Dynamic Neural Network Channel Execution for Efficient Training
- Leveraging synthetic imagery for collision-at-sea avoidance
- Disentangling Content and Style via Unsupervised Geometry Distillation
- Tuned Inception V3 for Recognizing States of Cooking Ingredients
- Fingerprint Spoof Buster
- Zero-Shot Learning via Category-Specific Visual-Semantic Mapping
- Deep Regression Forests for Age Estimation
- HeNet: A Deep Learning Approach on Intel Processor Trace for Effective Exploit Detection
- 3D-DETNet: a Single Stage Video-Based Vehicle Detector
- Class label autoencoder for zero-shot learning
- Fully Learnable Group Convolution for Acceleration of Deep Neural Networks
- FrameNet: Learning Local Canonical Frames of 3D Surfaces from a Single RGB Image
- Augmenting Gastrointestinal Health: A Deep Learning Approach to Human Stool Recognition and Characterization in Macroscopic Images
- Collaborative Layer-wise Discriminative Learning in Deep Neural Networks
- A Unified Multi-scale Deep Convolutional Neural Network for Fast Object Detection
- Artist Style Transfer Via Quadratic Potential
- SwiDeN : Convolutional Neural Networks For Depiction Invariant Object Recognition
- Labeling Topics with Images using Neural Networks
- Deep Learning for Low-Dose CT Denoising
- Quantifying contribution and propagation of error from computational steps, algorithms and hyperparameter choices in image classification pipelines
- 3D Robot Pose Estimation from 2D Images
- Mean Box Pooling: A Rich Image Representation and Output Embedding for the Visual Madlibs Task
- DeepDiary: Automatic Caption Generation for Lifelogging Image Streams
- Visual Question: Predicting If a Crowd Will Agree on the Answer
- Understanding the Impact of Label Granularity on CNN-based Image Classification
- Deep-learning-based identification of odontogenic keratocysts in hematoxylin- and eosin-stained jaw cyst specimens
- Learning Continuous Face Age Progression: A Pyramid of GANs
- Online Learning to Rank with List-level Feedback for Image Filtering
- Adaptive Fusion for RGB-D Salient Object Detection
- Edge-Semantic Learning Strategy for Layout Estimation in Indoor Environment
- Mixed context networks for semantic segmentation
- An Adaptive Approach for Automated Grapevine Phenotyping using VGG-based Convolutional Neural Networks
- Learning from Web Data: the Benefit of Unsupervised Object Localization
- Scene Labeling using Gated Recurrent Units with Explicit Long Range Conditioning
- Auto-tuning Neural Network Quantization Framework for Collaborative Inference Between the Cloud and Edge
- Beyond One Glance: Gated Recurrent Architecture for Hand Segmentation
- Multiple Discrimination and Pairwise CNN for View-based 3D Object Retrieval
- Optimal Gradient Quantization Condition for Communication-Efficient Distributed Training
- Focus on Semantic Consistency for Cross-domain Crowd Understanding
- Unsupervised Temporal Feature Aggregation for Event Detection in Unstructured Sports Videos
- Learning to Incorporate Structure Knowledge for Image Inpainting
- Low-rank Bilinear Pooling for Fine-Grained Classification
- Comprehensive and Efficient Data Labeling via Adaptive Model Scheduling
- Deep Multi-Modal Image Correspondence Learning
- Re-identification of Humans in Crowds using Personal, Social and Environmental Constraints
- Review: deep learning on 3D point clouds
- Generalized Deep Image to Image Regression
- On Iterative Neural Network Pruning, Reinitialization, and the Similarity of Masks
- Head and Tail Localization of C. elegans
- Neural ODEs for Image Segmentation with Level Sets
- ASAP: Asynchronous Approximate Data-Parallel Computation
- DeMIAN: Deep Modality Invariant Adversarial Network
- Online Knowledge Distillation with Diverse Peers
- EDAS: Efficient and Differentiable Architecture Search
- Machine: The New Art Connoisseur
- Variable Selection with Rigorous Uncertainty Quantification using Deep Bayesian Neural Networks: Posterior Concentration and Bernstein-von Mises Phenomenon
- Incremental Learning for Robot Perception through HRI
- Compression of Deep Neural Networks for Image Instance Retrieval
- AttKGCN: Attribute Knowledge Graph Convolutional Network for Person Re-identification
- Classification-driven Single Image Dehazing
- Optimal Mini-Batch Size Selection for Fast Gradient Descent
- Positional Attention-based Frame Identification with BERT: A Deep Learning Approach to Target Disambiguation and Semantic Frame Selection
- Random Forests and VGG-NET: An Algorithm for the ISIC 2017 Skin Lesion Classification Challenge
- Multi-Modal Machine Learning for Flood Detection in News, Social Media and Satellite Sequences
- Neural Language Priors
- MixModule: Mixed CNN Kernel Module for Medical Image Segmentation
- A Novel Self-Supervised Re-labeling Approach for Training with Noisy Labels
- Automatic Construction of Multi-layer Perceptron Network from Streaming Examples
- GraphX-Convolution for Point Cloud Deformation in 2D-to-3D Conversion
- Multi-level Wavelet Convolutional Neural Networks
- SADIH: Semantic-Aware DIscrete Hashing
- Modality-specific Cross-modal Similarity Measurement with Recurrent Attention Network
- Fixed smooth convolutional layer for avoiding checkerboard artifacts in CNNs
- Detection Method Based on Automatic Visual Shape Clustering for Pin-Missing Defect in Transmission Lines
- Ladder Loss for Coherent Visual-Semantic Embedding
- Im2Pencil: Controllable Pencil Illustration from Photographs
- TopoResNet: A hybrid deep learning architecture and its application to skin lesion classification
- Generic Tubelet Proposals for Action Localization
- Deep Learning with Energy-efficient Binary Gradient Cameras
- Viraliency: Pooling Local Virality
- Aff-Wild Database and AffWildNet
- Deep Fusion Network for Image Completion
- TailorGAN: Making User-Defined Fashion Designs
- Improved Deep Learning of Object Category using Pose Information
- Room Geometry Estimation from Room Impulse Responses using Convolutional Neural Networks
- 3D Object Classification via Spherical Projections
- Video Saliency Prediction Using Enhanced Spatiotemporal Alignment Network
- A Joint Model for Multimodal Document Quality Assessment
- Butterfly Effect: Bidirectional Control of Classification Performance by Small Additive Perturbation
- Learning Temporal Embeddings for Complex Video Analysis
- A novel method for identifying the deep neural network model with the Serial Number
- Correlation Congruence for Knowledge Distillation
- Obstruction level detection of sewer videos using convolutional neural networks
- Geometry of Deep Convolutional Networks
- To Boost or Not to Boost? On the Limits of Boosted Trees for Object Detection
- Deep learning-based phase control method for coherent beam combining and its application in generating orbital angular momentum beams
- ID-aware Quality for Set-based Person Re-identification
- A Robust and Precise ConvNet for small non-coding RNA classification (RPC-snRC)
- Deleting object selective units in a fully-connected layer of deep convolutional networks improves classification performance
- Multi-view PointNet for 3D Scene Understanding
- Global Adversarial Attacks for Assessing Deep Learning Robustness
- SpotTheFake: An Initial Report on a New CNN-Enhanced Platform for Counterfeit Goods Detection
- Shared Mobile-Cloud Inference for Collaborative Intelligence
- Lexicon-Free Fingerspelling Recognition from Video: Data, Models, and Signer Adaptation
- Visual Semantic Information Pursuit: A Survey
- Discoverability in Satellite Imagery: A Good Sentence is Worth a Thousand Pictures
- Being the center of attention: A Person-Context CNN framework for Personality Recognition
- Effective Image Retrieval via Multilinear Multi-index Fusion
- Efficient Neural Task Adaptation by Maximum Entropy Initialization
- Kinship Verification from Videos using Spatio-Temporal Texture Features and Deep Learning
- Rethinking the Artificial Neural Networks: A Mesh of Subnets with a Central Mechanism for Storing and Predicting the Data
- Neurlux: Dynamic Malware Analysis Without Feature Engineering
- A Taught-Obesrve-Ask (TOA) Method for Object Detection with Critical Supervision
- Efficient Cloth Simulation using Miniature Cloth and Upscaling Deep Neural Networks
- Exploiting Operation Importance for Differentiable Neural Architecture Search
- Enhancing Salient Object Segmentation Through Attention
- Evolutionary Deep Learning to Identify Galaxies in the Zone of Avoidance
- Metric Classification Network in Actual Face Recognition Scene
- Visual aesthetic analysis using deep neural network: model and techniques to increase accuracy without transfer learning
- Learning Manipulation under Physics Constraints with Visual Perception
- Reference Product Search
- Real Time Fine-Grained Categorization with Accuracy and Interpretability
- Brain-inspired reverse adversarial examples
- Deep Deformable Registration: Enhancing Accuracy by Fully Convolutional Neural Net
- A Methodological Review of Visual Road Recognition Procedures for Autonomous Driving Applications
- swCaffe: a Parallel Framework for Accelerating Deep Learning Applications on Sunway TaihuLight
- Fully Convolutional Neural Network for Semantic Segmentation of Anatomical Structure and Pathologies in Colour Fundus Images Associated with Diabetic Retinopathy
- LanCe: A Comprehensive and Lightweight CNN Defense Methodology against Physical Adversarial Attacks on Embedded Multimedia Applications
- 3D Face Mask Presentation Attack Detection Based on Intrinsic Image Analysis
- Ancient Painting to Natural Image: A New Solution for Painting Processing
- Large-scale image analysis using docker sandboxing
- Deep Visual City Recognition Visualization
- Transfer learning from language models to image caption generators: Better models may not transfer better
- Generative Collaborative Networks for Single Image Super-Resolution
- Object-Level Context Modeling For Scene Classification with Context-CNN
- A Generative Model for Sampling High-Performance and Diverse Weights for Neural Networks
- Generalization Tower Network: A Novel Deep Neural Network Architecture for Multi-Task Learning
- Attention Transfer from Web Images for Video Recognition
- Bandlimiting Neural Networks Against Adversarial Attacks
- ResNetX: a more disordered and deeper network architecture
- Deeply-Supervised CNN for Prostate Segmentation
- Human-like machine thinking: Language guided imagination
- Maxmin convolutional neural networks for image classification
- Learning Concept Taxonomies from Multi-modal Data
- Efficient Training of Convolutional Neural Nets on Large Distributed Systems
- Example-Guided Scene Image Synthesis using Masked Spatial-Channel Attention and Patch-Based Self-Supervision
- Rethinking Convolutional Semantic Segmentation Learning
- WeNet: Weighted Networks for Recurrent Network Architecture Search
- Data-Driven Compression of Convolutional Neural Networks
- Pointwise Attention-Based Atrous Convolutional Neural Networks
- Deep Neural Network Assisted Iterative Reconstruction Method for Low Dose CT
- NeuralVis: Visualizing and Interpreting Deep Learning Models
- Visual Weather Temperature Prediction
- Improving training of deep neural networks via Singular Value Bounding
- Controllable List-wise Ranking for Universal No-reference Image Quality Assessment
- Neural Style Transfer for Point Clouds
- Binarized Neural Architecture Search
- Scheduled Differentiable Architecture Search for Visual Recognition
- Domain Randomization for Active Pose Estimation
- Automated Segmentation of the Optic Disk and Cup using Dual-Stage Fully Convolutional Networks
- Differential Recurrent Neural Network and its Application for Human Activity Recognition
- Very Long Natural Scenery Image Prediction by Outpainting
- Deep learning at scale for subgrid modeling in turbulent flows
- Automatic Data Augmentation by Learning the Deterministic Policy
- Using KL-divergence to focus Deep Visual Explanation
- Patchy Image Structure Classification Using Multi-Orientation Region Transform
- Water Preservation in Soan River Basin using Deep Learning Techniques
- Towards the Next Generation of Retinal Neuroprosthesis: Visual Computation with Spikes
- L3 Fusion: Fast Transformed Convolutions on CPUs
- Annotation and Detection of Emotion in Text-based Dialogue Systems with CNN
- Where is my Phone ? Personal Object Retrieval from Egocentric Images
- A Deep Decoder Structure Based on WordEmbedding Regression for An Encoder-Decoder Based Model for Image Captioning
- Maximal adversarial perturbations for obfuscation: Hiding certain attributes while preserving rest
- StyleNAS: An Empirical Study of Neural Architecture Search to Uncover Surprisingly Fast End-to-End Universal Style Transfer Networks
- Controlling biases and diversity in diverse image-to-image translation
- Predicting Visual Memory Schemas with Variational Autoencoders
- Winter Road Surface Condition Recognition Using A Pretrained Deep Convolutional Network
- Patch Aggregator for Scene Text Script Identification
- End-to-End Time-Lapse Video Synthesis from a Single Outdoor Image
- Regularize, Expand and Compress: Multi-task based Lifelong Learning via NonExpansive AutoML
- Incorporating Temporal Prior from Motion Flow for Instrument Segmentation in Minimally Invasive Surgery Video
- A Paradigm Shift: Detecting Human Rights Violations Through Web Images
- HGC: Hierarchical Group Convolution for Highly Efficient Neural Network
- OGNet: Salient Object Detection with Output-guided Attention Module
- Dually Supervised Feature Pyramid for Object Detection and Segmentation
- Deep Learning for Identifying Potential Conceptual Shifts for Co-creative Drawing
- Pruning Convolutional Neural Networks for Image Instance Retrieval
- AutoScaler: Scale-Attention Networks for Visual Correspondence
- RAPDARTS: Resource-Aware Progressive Differentiable Architecture Search
- Unsupervised Temperature Scaling: An Unsupervised Post-Processing Calibration Method of Deep Networks
- Less-forgetful Learning for Domain Expansion in Deep Neural Networks
- Spatial-Aware Non-Local Attention for Fashion Landmark Detection
- Logo-2K+: A Large-Scale Logo Dataset for Scalable Logo Classification
- Sensitivity of Deep Convolutional Networks to Gabor Noise
- On the Learning Property of Logistic and Softmax Losses for Deep Neural Networks
- EXPERTNet Exigent Features Preservative Network for Facial Expression Recognition
- Fine-grained Uncertainty Modeling in Neural Networks
- Recognizing Video Events with Varying Rhythms
- Joint Dictionaries for Zero-Shot Learning
- Learning with Out-of-Distribution Data for Audio Classification
- Material Classification using Neural Networks
- EPNAS: Efficient Progressive Neural Architecture Search
- Few-Features Attack to Fool Machine Learning Models through Mask-Based GAN
- Learning Orientation-Estimation Convolutional Neural Network for Building Detection in Optical Remote Sensing Image
- Unifying Heterogeneous Classifiers with Distillation
- Context Augmentation for Convolutional Neural Networks
- Locality Preserving Joint Transfer for Domain Adaptation
- Short-Term Temporal Convolutional Networks for Dynamic Hand Gesture Recognition
- Hand Orientation Estimation in Probability Density Form
- Deep Learning in Memristive Nanowire Networks
- Detecting Patch Adversarial Attacks with Image Residuals
- Latent Adversarial Defence with Boundary-guided Generation
- Practical License Plate Recognition in Unconstrained Surveillance Systems with Adversarial Super-Resolution
- Attention Monitoring and Hazard Assessment with Bio-Sensing and Vision: Empirical Analysis Utilizing CNNs on the KITTI Dataset
- ZoomCount: A Zooming Mechanism for Crowd Counting in Static Images
- 5D Light Field Synthesis from a Monocular Video
- Neuron ranking -- an informed way to condense convolutional neural networks architecture
- EgoReID: Cross-view Self-Identification and Human Re-identification in Egocentric and Surveillance Videos
- AuxBlocks: Defense Adversarial Example via Auxiliary Blocks
- Incomplete Dot Products for Dynamic Computation Scaling in Neural Network Inference
- Stagioni: Temperature management to enable near-sensor processing for energy-efficient high-fidelity imaging
- PoseConvGRU: A Monocular Approach for Visual Ego-motion Estimation by Learning
- A simple and effective postprocessing method for image classification
- Distill-2MD-MTL: Data Distillation based on Multi-Dataset Multi-Domain Multi-Task Frame Work to Solve Face Related Tasksks, Multi Task Learning, Semi-Supervised Learning
- An Overview on Data Representation Learning: From Traditional Feature Learning to Recent Deep Learning
- Semi-supervised Learning on Graph with an Alternating Diffusion Process
- The VQA-Machine: Learning How to Use Existing Vision Algorithms to Answer New Questions
- An Integrated Approach to Crowd Video Analysis: From Tracking to Multi-level Activity Recognition
- A Locating Model for Pulmonary Tuberculosis Diagnosis in Radiographs
- Sparse Factorization Layers for Neural Networks with Limited Supervision
- Functional Error Correction for Robust Neural Networks
- Detect, Replace, Refine: Deep Structured Prediction For Pixel Wise Labeling
- Learning Filter Banks Using Deep Learning For Acoustic Signals
- Efficient Inferencing of Compressed Deep Neural Networks
- Convolutional Dictionary Pair Learning Network for Image Representation Learning
- Automatic Lyrics Alignment and Transcription in Polyphonic Music: Does Background Music Help?
- Impact of ImageNet Model Selection on Domain Adaptation
- Is a Picture Worth Ten Thousand Words in a Review Dataset?
- Fine-Grained Urban Flow Inference
- SalGaze: Personalizing Gaze Estimation Using Visual Saliency
- Towards Structured Evaluation of Deep Neural Network Supervisors
- Large Batch Training Does Not Need Warmup
- Clothing Retrieval with Visual Attention Model
- Improving the Evaluation of Generative Models with Fuzzy Logic
- 3D Shape Segmentation with Geometric Deep Learning
- Hardware-Driven Nonlinear Activation for Stochastic Computing Based Deep Convolutional Neural Networks
- Joint Concept Matching based Learning for Zero-Shot Recognition
- Drone Path-Following in GPS-Denied Environments using Convolutional Networks
- On the Role of Receptive Field in Unsupervised Sim-to-Real Image Translation
- Diverse Sampling for Self-Supervised Learning of Semantic Segmentation
- Reasoning about Fine-grained Attribute Phrases using Reference Games
- Yelp Food Identification via Image Feature Extraction and Classification
- Learning Dynamic Hierarchical Models for Anytime Scene Labeling
- Projectron -- A Shallow and Interpretable Network for Classifying Medical Images
- Deep CSI Learning for Gait Biometric Sensing and Recognition
- High Frequency Residual Learning for Multi-Scale Image Classification
- Communication Lower Bound in Convolution Accelerators
- Representation of White- and Black-Box Adversarial Examples in Deep Neural Networks and Humans: A Functional Magnetic Resonance Imaging Study
- Counting Cells in Time-Lapse Microscopy using Deep Neural Networks
- Localization-aware Channel Pruning for Object Detection
- UNO: Uncertainty-aware Noisy-Or Multimodal Fusion for Unanticipated Input Degradation
- Where and Who? Automatic Semantic-Aware Person Composition
- Human Face Expressions from Images - 2D Face Geometry and 3D Face Local Motion versus Deep Neural Features
- A Selfie is Worth a Thousand Words: Mining Personal Patterns behind User Selfie-posting Behaviours
- Self-Supervised Visual Representations for Cross-Modal Retrieval
- Transfer Learning in 4D for Breast Cancer Diagnosis using Dynamic Contrast-Enhanced Magnetic Resonance Imaging
- The Pitfall of Evaluating Performance on Emerging AI Accelerators
- An unsupervised long short-term memory neural network for event detection in cell videos
- Ship classification from overhead imagery using synthetic data and domain adaptation
- Grouping Capsules Based Different Types
- Real-time Multiple People Hand Localization in 4D Point Clouds
- American Sign Language fingerspelling recognition from video: Methods for unrestricted recognition and signer-independence
- Improving Image Captioning by Leveraging Knowledge Graphs
- Optimistic and Pessimistic Neural Networks for Scene and Object Recognition
- Learning Local Shape Descriptors from Part Correspondences With Multi-view Convolutional Networks
- CartoonRenderer: An Instance-based Multi-Style Cartoon Image Translator
- Modeling Image Virality with Pairwise Spatial Transformer Networks
- EvaluationNet: Can Human Skill be Evaluated by Deep Networks?
- Discriminatively Learned Hierarchical Rank Pooling Networks
- CODA: Counting Objects via Scale-aware Adversarial Density Adaption
- Examining Representational Similarity in ConvNets and the Primate Visual Cortex
- LookUP: Vision-Only Real-Time Precise Underground Localisation for Autonomous Mining Vehicles
- A CNN-RNN Architecture for Multi-Label Weather Recognition
- Recognizing Dynamic Scenes with Deep Dual Descriptor based on Key Frames and Key Segments
- Learning to Forecast Videos of Human Activity with Multi-granularity Models and Adaptive Rendering
- Automatic Handgun Detection Alarm in Videos Using Deep Learning
- Semantic Segmentation of Earth Observation Data Using Multimodal and Multi-scale Deep Networks
- Automatic Lumbar Spinal CT Image Segmentation with a Dual Densely Connected U-Net
- Ensemble of Part Detectors for Simultaneous Classification and Localization
- A witness function based construction of discriminative models using Hermite polynomials
- Energy Saving Additive Neural Network
- Balanced Quantization: An Effective and Efficient Approach to Quantized Neural Networks
- StyleRemix: An Interpretable Representation for Neural Image Style Transfer
- DARB: A Density-Aware Regular-Block Pruning for Deep Neural Networks
- Balancing Specialization, Generalization, and Compression for Detection and Tracking
- Air Quality Measurement Based on Double-Channel Convolutional Neural Network Ensemble Learning
- ARCHANGEL: Tamper-proofing Video Archives using Temporal Content Hashes on the Blockchain
- Enhancing Sound Texture in CNN-Based Acoustic Scene Classification
- HAL: Improved Text-Image Matching by Mitigating Visual Semantic Hubs
- Spotting insects from satellites: modeling the presence of Culicoides imicola through Deep CNNs
- Deceiving Google's Cloud Video Intelligence API Built for Summarizing Videos
- An approach to image denoising using manifold approximation without clean images
- Learning Competitive and Discriminative Reconstructions for Anomaly Detection
- Tutorial on Answering Questions about Images with Deep Learning
- Improving image generative models with human interactions
- Visual Dialogue State Tracking for Question Generation
- Learning Shared Semantic Space with Correlation Alignment for Cross-modal Event Retrieval
- WSOD with PSNet and Box Regression
- CrescendoNet: A Simple Deep Convolutional Neural Network with Ensemble Behavior
- CP-decomposition with Tensor Power Method for Convolutional Neural Networks Compression
- Fast Training of Convolutional Neural Networks via Kernel Rescaling
- Large Scale Incremental Learning
- Master's Thesis : Deep Learning for Visual Recognition
- Multimodal Deep Network Embedding with Integrated Structure and Attribute Information
- Hardware-friendly Neural Network Architecture for Neuromorphic Computing
- Few-Shot Abstract Visual Reasoning With Spectral Features
- Parametrization of stochastic inputs using generative adversarial networks with application in geology
- Bringing Impressionism to Life with Neural Style Transfer in Come Swim
- Temporal Hockey Action Recognition via Pose and Optical Flows
- Dual Asymmetric Deep Hashing Learning
- Deep Object Co-segmentation via Spatial-Semantic Network Modulation
- Quantifying error contributions of computational steps, algorithms and hyperparameter choices in image classification pipelines
- Design of a Very Compact CNN Classifier for Online Handwritten Chinese Character Recognition Using DropWeight and Global Pooling
- Revisiting IM2GPS in the Deep Learning Era
- Typed Graph Networks
- Detecting cutaneous basal cell carcinomas in ultra-high resolution and weakly labelled histopathological images
- Learning On-Road Visual Control for Self-Driving Vehicles with Auxiliary Tasks
- Compositional Temporal Visual Grounding of Natural Language Event Descriptions
- Detecting Kissing Scenes in a Database of Hollywood Films
- Delving Deep into Liver Focal Lesion Detection: A Preliminary Study
- Fine-grained Classification of Rowing teams
- Radial Prediction Layer
- Joint Viewpoint and Keypoint Estimation with Real and Synthetic Data
- Improved Reinforcement Learning through Imitation Learning Pretraining Towards Image-based Autonomous Driving
- Neural Networks Weights Quantization: Target None-retraining Ternary (TNT)
- Automated Weed Detection in Aerial Imagery with Context
- Inserting Videos into Videos
- Learning to Recognize Objects by Retaining other Factors of Variation
- ProductNet: a Collection of High-Quality Datasets for Product Representation Learning
- A Sketch Based 3D Shape Retrieval Approach Based on Efficient Deep Point-to-Subspace Metric Learning
- A Capsule-unified Framework of Deep Neural Networks for Graphical Programming
- Person Identification with Visual Summary for a Safe Access to a Smart Home
- ESFNet: Efficient Network for Building Extraction from High-Resolution Aerial Images
- Weakly supervised object detection using pseudo-strong labels
- Automatic Attribute Discovery with Neural Activations
- Latent Variable Algorithms for Multimodal Learning and Sensor Fusion
- Learning Actor Relation Graphs for Group Activity Recognition
- Face Image Reflection Removal
- Quaternion Convolutional Neural Networks
- A General Framework for Edited Video and Raw Video Summarization
- A data-driven proxy to Stoke's flow in porous media
- FaceLiveNet+: A Holistic Networks For Face Authentication Based On Dynamic Multi-task Convolutional Neural Networks
- CircConv: A Structured Convolution with Low Complexity
- To believe or not to believe: Validating explanation fidelity for dynamic malware analysis
- Transparency guided ensemble convolutional neural networks for stratification of pseudoprogression and true progression of glioblastoma multiform
- Automated X-ray Image Analysis for Cargo Security: Critical Review and Future Promise
- A Recursive Framework for Expression Recognition: From Web Images to Deep Models to Game Dataset
- Min-Entropy Latent Model for Weakly Supervised Object Detection
- Multi-Mode Inference Engine for Convolutional Neural Networks
- Box-level Segmentation Supervised Deep Neural Networks for Accurate and Real-time Multispectral Pedestrian Detection
- Towards thinner convolutional neural networks through Gradually Global Pruning
- Interest-Related Item Similarity Model Based on Multimodal Data for Top-N Recommendation
- Manifestation of Image Contrast in Deep Networks
- Simulating CRF with CNN for CNN
- Every Filter Extracts A Specific Texture In Convolutional Neural Networks
- Dynamic Multi-path Neural Network
- Explicit topological priors for deep-learning based image segmentation using persistent homology
- Structural Material Property Tailoring Using Deep Neural Networks
- ICLR Reproducibility Challenge Report (Padam : Closing The Generalization Gap Of Adaptive Gradient Methods in Training Deep Neural Networks)
- Information-Theoretic Understanding of Population Risk Improvement with Model Compression
- Classifier comparison using precision
- A scalable convolutional neural network for task-specified scenarios via knowledge distillation
- One-Shot Image-to-Image Translation via Part-Global Learning with a Multi-adversarial Framework
- Learning with Collaborative Neural Network Group by Reflection
- Label distribution based facial attractiveness computation by deep residual learning
- VGG Fine-tuning for Cooking State Recognition
- Implicit Filter Sparsification In Convolutional Neural Networks
- "Tom" pet robot applied to urban autism
- Light-weighted Saliency Detection with Distinctively Lower Memory Cost and Model Size
- Watermark retrieval from 3D printed objects via synthetic data training
- Automatic Detection of Knee Joints and Quantification of Knee Osteoarthritis Severity using Convolutional Neural Networks
- Pose estimator and tracker using temporal flow maps for limbs
- Scalable Discrete Supervised Hash Learning with Asymmetric Matrix Factorization
- Depth Map Completion by Jointly Exploiting Blurry Color Images and Sparse Depth Maps
- Context-modulation of hippocampal dynamics and deep convolutional networks
- Towards Efficient Neural Networks On-a-chip: Joint Hardware-Algorithm Approaches
- StuffNet: Using 'Stuff' to Improve Object Detection
- AccUDNN: A GPU Memory Efficient Accelerator for Training Ultra-deep Neural Networks
- Towards end-to-end optimisation of functional image analysis pipelines
- Vehicle Detection in Deep Learning
- Toward Runtime-Throttleable Neural Networks
- Orthogonal and Idempotent Transformations for Learning Deep Neural Networks
- CamLoc: Pedestrian Location Detection from Pose Estimation on Resource-constrained Smart-cameras
- DE-PACRR: Exploring Layers Inside the PACRR Model
- Sampling Using Neural Networks for colorizing the grayscale images
- Studying the Plasticity in Deep Convolutional Neural Networks using Random Pruning
- Classification of X-Ray Protein Crystallization Using Deep Convolutional Neural Networks with a Finder Module
- Neural networks versus Logistic regression for 30 days all-cause readmission prediction
- Do Deep Neural Networks Suffer from Crowding?
- Voronoi-based compact image descriptors: Efficient Region-of-Interest retrieval with VLAD and deep-learning-based descriptors
- Recognizing and Presenting the Storytelling Video Structure with Deep Multimodal Networks
- Parallel Attention: A Unified Framework for Visual Object Discovery through Dialogs and Queries
- Progressive Cluster Purification for Transductive Few-shot Learning
- Efficient Convolutional Neural Network with Binary Quantization Layer
- Low Photon Budget Phase Retrieval with Perceptual Loss Trained Deep Neural Networks
- Learning from Positive and Unlabeled Data by Identifying the Annotation Process
- Know What You Don't Know: Modeling a Pragmatic Speaker that Refers to Objects of Unknown Categories
- A Deep Framework for Bone Age Assessment based on Finger Joint Localization
- Neural Inheritance Relation Guided One-Shot Layer Assignment Search
- Understanding Natural Language Instructions for Fetching Daily Objects Using GAN-Based Multimodal Target-Source Classification
- A Self Validation Network for Object-Level Human Attention Estimation
- Assembling Semantically-Disentangled Representations for Predictive-Generative Models via Adaptation from Synthetic Domain
- Improvement of Multiparametric MR Image Segmentation by Augmenting the Data with Generative Adversarial Networks for Glioma Patients
- A Molecular-MNIST Dataset for Machine Learning Study on Diffraction Imaging and Microscopy
- DeepDualMapper: A Gated Fusion Network for Automatic Map Extraction using Aerial Images and Trajectories
- Facial Attribute Capsules for Noise Face Super Resolution
- CRL: Class Representative Learning for Image Classification
- Efficient Multi-Domain Network Learning by Covariance Normalization
- Saliency Detection With Fully Convolutional Neural Network
- Understanding and Predicting The Attractiveness of Human Action Shot
- Multi-scale recognition with DAG-CNNs
- DeepTEGINN: Deep Learning Based Tools to Extract Graphs from Images of Neural Networks
- Deep-Geometric 6 DoF Localization from a Single Image in Topo-metric Maps
- Region-Manipulated Fusion Networks for Pancreatitis Recognition
- DVNet: A Memory-Efficient Three-Dimensional CNN for Large-Scale Neurovascular Reconstruction
- Large Hole Image Inpainting With Compress-Decompression Network
- Learning Deep Analysis Dictionaries -- Part II: Convolutional Dictionaries
- Learning to Catch Piglets in Flight
- AE-OT-GAN: Training GANs from data specific latent distribution
- Brain Metastasis Segmentation Network Trained with Robustness to Annotations with Multiple False Negatives
- Point-of-Care Diabetic Retinopathy Diagnosis: A Standalone Mobile Application Approach
- RFBTD: RFB Text Detector
- Image Patch Matching Using Convolutional Descriptors with Euclidean Distance
- Adaptive Loss Function for Super Resolution Neural Networks Using Convex Optimization Techniques
- A Comprehensive Study on Temporal Modeling for Online Action Detection
- Predicting Ground-Level Scene Layout from Aerial Imagery
- Segmentation-by-Detection: A Cascade Network for Volumetric Medical Image Segmentation
- Real-Time Lane ID Estimation Using Recurrent Neural Networks With Dual Convention
- Weighted parallel SGD for distributed unbalanced-workload training system
- Beam Search for Learning a Deep Convolutional Neural Network of 3D Shapes
- Learning with Rethinking: Recurrently Improving Convolutional Neural Networks through Feedback
- Adaptive Weighting Depth-variant Deconvolution of Fluorescence Microscopy Images with Convolutional Neural Network
- Attending to Emotional Narratives
- Attention-Aware Answers of the Crowd
- Privacy for Rescue: A New Testimony Why Privacy is Vulnerable In Deep Models
- Continuous Speech Recognition using EEG and Video
- An efficient deep learning hashing neural network for mobile visual search
- One-Shot Fine-Grained Instance Retrieval
- Anticipating Daily Intention using On-Wrist Motion Triggered Sensing
- TOCO: A Framework for Compressing Neural Network Models Based on Tolerance Analysis
- NeuroTrainer: An Intelligent Memory Module for Deep Learning Training
- Deep Exemplar Networks for VQA and VQG
- Mitigate Parasitic Resistance in Resistive Crossbar-based Convolutional Neural Networks
- Rapid Identification of X-ray Diffraction Spectra Based on Very Limited Data by Interpretable Convolutional Neural Networks
- Progressive Representation Adaptation for Weakly Supervised Object Localization
- Deep inspection: an electrical distribution pole parts study via deep neural networks
- A Method for Arbitrary Instance Style Transfer
- CCCNet: An Attention Based Deep Learning Framework for Categorized Crowd Counting
- Arithmetic addition of two integers by deep image classification networks: experiments to quantify their autonomous reasoning ability
- Unified Signal Compression Using Generative Adversarial Networks
- An evaluation of large-scale methods for image instance and class discovery
- TasselNet: Counting maize tassels in the wild via local counts regression network
- Exploiting Motion Information from Unlabeled Videos for Static Image Action Recognition
- A Reconfigurable Streaming Deep Convolutional Neural Network Accelerator for Internet of Things
- ADA: A Game-Theoretic Perspective on Data Augmentation for Object Detection
- Chained Multi-stream Networks Exploiting Pose, Motion, and Appearance for Action Classification and Detection
- In-field grape berries counting for yield estimation using dilated CNNs
- Dynamical System Inspired Adaptive Time Stepping Controller for Residual Network Families
- Automating Image Analysis by Annotating Landmarks with Deep Neural Networks
- Segmenting Medical MRI via Recurrent Decoding Cell
- Dual Reconstruction with Densely Connected Residual Network for Single Image Super-Resolution
- Collaborative Deep Reinforcement Learning for Joint Object Search
- Deep learning-based assessment of tumor-associated stroma for diagnosing breast cancer in histopathology images
- Sensing Urban Land-Use Patterns By Integrating Google Tensorflow And Scene-Classification Models
- NGEMM: Optimizing GEMM for Deep Learning via Compiler-based Techniques
- Automatic Mouse Embryo Brain Ventricle & Body Segmentation and Mutant Classification From Ultrasound Data Using Deep Learning
- weedNet: Dense Semantic Weed Classification Using Multispectral Images and MAV for Smart Farming
- On Architectures for Including Visual Information in Neural Language Models for Image Description
- Enabling Highly Efficient Capsule Networks Processing Through A PIM-Based Architecture Design
- End to end collision avoidance based on optical flow and neural networks
- DeepBlindness: Fast Blindness Map Estimation and Blindness Type Classification for Outdoor Scene from Single Color Image
- Hashing in the Zero Shot Framework with Domain Adaptation
- Transfer Learning in Visual and Relational Reasoning
- Fast-UAP: An Algorithm for Speeding up Universal Adversarial Perturbation Generation with Orientation of Perturbation Vectors
- Fitness Done Right: a Real-time Intelligent Personal Trainer for Exercise Correction
- Integrated Deep and Shallow Networks for Salient Object Detection
- A Dataset for Developing and Benchmarking Active Vision
- Deep Collaborative Learning for Visual Recognition
- Spectral Algorithm for Low-rank Multitask Regression
- Compression Fractures Detection on CT
- Deep Classification Network for Monocular Depth Estimation
- Streaming Networks: Enable A Robust Classification of Noise-Corrupted Images
- Generative-Discriminative Variational Model for Visual Recognition
- CATERPILLAR: Coarse Grain Reconfigurable Architecture for Accelerating the Training of Deep Neural Networks
- Adaptively Denoising Proposal Collection for Weakly Supervised Object Localization
- Modelling the Scene Dependent Imaging in Cameras with a Deep Neural Network
- Joint Max Margin and Semantic Features for Continuous Event Detection in Complex Scenes
- Deep Learning for Prostate Pathology
- Beyond Monte Carlo Tree Search: Playing Go with Deep Alternative Neural Network and Long-Term Evaluation
- Multi-Path Region-Based Convolutional Neural Network for Accurate Detection of Unconstrained "Hard Faces"
- A Robust Indoor Scene Recognition Method based on Sparse Representation
- Tell-the-difference: Fine-grained Visual Descriptor via a Discriminating Referee
- EDEN: Enabling Energy-Efficient, High-Performance Deep Neural Network Inference Using Approximate DRAM
- Self-enhancement of automatic tunnel accident detection (TAD) on CCTV by AI deep-learning
- A Fully Trainable Network with RNN-based Pooling
- Cribriform pattern detection in prostate histopathological images using deep learning models
- Neural Models of the Psychosemantics of `Most'
- Semantic Segmentation via Highly Fused Convolutional Network with Multiple Soft Cost Functions
- EnKCF: Ensemble of Kernelized Correlation Filters for High-Speed Object Tracking
- Novel digital tissue phenotypic signatures of distant metastasis in colorectal cancer
- Multi-vision Attention Networks for On-line Red Jujube Grading
- Dance Dance Generation: Motion Transfer for Internet Videos
- Weakly Supervised Object Detection with Pointwise Mutual Information
- Deep Neural Networks In Fully Connected CRF For Image Labeling With Social Network Metadata
- DeepSIC: Deep Semantic Image Compression
- Multi-scale discriminative Region Discovery for Weakly-Supervised Object Localization
- Classification of sparsely labeled spatio-temporal data through semi-supervised adversarial learning
- Rhythmic Representations: Learning Periodic Patterns for Scalable Place Recognition at a Sub-Linear Storage Cost
- High-Level Perceptual Similarity is Enabled by Learning Diverse Tasks
- Macrocanonical Models for Texture Synthesis
- Zero-Shot Learning with Multi-Battery Factor Analysis
- Deep Octonion Networks
- Signed Input Regularization
- Depth Assisted Full Resolution Network for Single Image-based View Synthesis
- News Cover Assessment via Multi-task Learning
- Pseudo-positive regularization for deep person re-identification
- Sentiment Analysis from Images of Natural Disasters
- Multi Instance Learning For Unbalanced Data
- Training Deep Neural Networks to Detect Repeatable 2D Features Using Large Amounts of 3D World Capture Data
- Fine-Grained Categorization via CNN-Based Automatic Extraction and Integration of Object-Level and Part-Level Features
- Vision Recognition using Discriminant Sparse Optimization Learning
- DaiMoN: A Decentralized Artificial Intelligence Model Network
- A DNN Framework For Text Image Rectification From Planar Transformations
- Automatic discovery of discriminative parts as a quadratic assignment problem
- Non-uniqueness phenomenon of object representation in modelling IT cortex by deep convolutional neural network (DCNN)
- Understanding Adversarial Behavior of DNNs by Disentangling Non-Robust and Robust Components in Performance Metric
- Online Hashing with Efficient Updating of Binary Codes
- Adaptive ROI Generation for Video Object Segmentation Using Reinforcement Learning
- A backward pass through a CNN using a generative model of its activations
- Kill Two Birds with One Stone: Weakly-Supervised Neural Network for Image Annotation and Tag Refinement
- Image2song: Song Retrieval via Bridging Image Content and Lyric Words
- A Deep Convolutional Network for Seismic Shot-Gather Image Quality Classification
- Prune the Convolutional Neural Networks with Sparse Shrink
- LocalNorm: Robust Image Classification through Dynamically Regularized Normalization
- Corn leaf detection using Region based convolutional neural network
- Computational Anatomy for Multi-Organ Analysis in Medical Imaging: A Review
- Memeify: A Large-Scale Meme Generation System
- Using Satellite Imagery for Good: Detecting Communities in Desert and Mapping Vaccination Activities
- An Empirical Analysis of Deep Audio-Visual Models for Speech Recognition
- Unsupervised Projection Networks for Generative Adversarial Networks
- Canonical Correlation Analysis for Misaligned Satellite Image Change Detection
- Cascaded Coarse-to-Fine Deep Kernel Networks for Efficient Satellite Image Change Detection
- Block-Cyclic Stochastic Coordinate Descent for Deep Neural Networks
- Sparse Signal Recovery for Binary Compressed Sensing by Majority Voting Neural Networks
- A Generalization Theory based on Independent and Task-Identically Distributed Assumption
- Deep Policy Hashing Network with Listwise Supervision
- Patch Reordering: a Novel Way to Achieve Rotation and Translation Invariance in Convolutional Neural Networks
- Exploring the Challenges towards Lifelong Fact Learning
- Anatomical labeling of brain CT scan anomalies using multi-context nearest neighbor relation networks
- LucidDream: Controlled Temporally-Consistent DeepDream on Videos
- GRIm-RePR: Prioritising Generating Important Features for Pseudo-Rehearsal
- Weakly-Supervised Road Affordances Inference and Learning in Scenes without Traffic Signs
- Decision Propagation Networks for Image Classification
- Diversity Promoting Online Sampling for Streaming Video Summarization
- Unsupervised monocular stereo matching
- Closed-Loop Adaptation for Weakly-Supervised Semantic Segmentation
- Multiple Instance Learning Convolutional Neural Networks for Object Recognition
- Semantic Softmax Loss for Zero-Shot Learning
- Occluded Pedestrian Detection with Visible IoU and Box Sign Predictor
- Super Interaction Neural Network
- Speeding up convolutional networks pruning with coarse ranking
- User-centric Composable Services: A New Generation of Personal Data Analytics
- Generating Adversarial Perturbation with Root Mean Square Gradient
- Cascaded Detail-Preserving Networks for Super-Resolution of Document Images
- Radius Adaptive Convolutional Neural Network
- A Novel Unsupervised Post-Processing Calibration Method for DNNS with Robustness to Domain Shift
- Deep image representations using caption generators
- Local Area Transform for Cross-Modality Correspondence Matching and Deep Scene Recognition
- Semi-supervised Fisher vector network
- Instance Map based Image Synthesis with a Denoising Generative Adversarial Network
- FusionStitching: Boosting Execution Efficiency of Memory Intensive Computations for DL Workloads
- Vision-based deep execution monitoring
- C2S2: Cost-aware Channel Sparse Selection for Progressive Network Pruning
- Design, Analysis and Application of A Volumetric Convolutional Neural Network
- Locality Constraint Dictionary Learning with Support Vector for Pattern Classification
- Scalable Object Detection for Stylized Objects
- Coupled Support Vector Machines for Supervised Domain Adaptation
- A CNN Cascade for Landmark Guided Semantic Part Segmentation
- Supervised Online Hashing via Hadamard Codebook Learning
- How Compact?: Assessing Compactness of Representations through Layer-Wise Pruning
- Depth Not Needed - An Evaluation of RGB-D Feature Encodings for Off-Road Scene Understanding by Convolutional Neural Network
- Video Segment Copy Detection Using Memory Constrained Hierarchical Batch-Normalized LSTM Autoencoder
- Low-Cost Transfer Learning of Face Tasks
- BitSplit-Net: Multi-bit Deep Neural Network with Bitwise Activation Function
- Latent Hinge-Minimax Risk Minimization for Inference from a Small Number of Training Samples
- Inspect Transfer Learning Architecture with Dilated Convolution
- Reinforced Bit Allocation under Task-Driven Semantic Distortion Metrics
- Shift Convolution Network for Stereo Matching
- Spatial Sampling Network for Fast Scene Understanding
- Slim-DP: A Light Communication Data Parallelism for DNN
- Learning to Label Affordances from Simulated and Real Data
- Optimizing Memory Efficiency for Convolution Kernels on Kepler GPUs
- Lecture video indexing using boosted margin maximizing neural networks
- Generating Minimal Adversarial Perturbations with Integrated Adaptive Gradients
- Image retrieval method based on CNN and dimension reduction
- Deep CTR Prediction in Display Advertising
- FRAME Revisited: An Interpretation View Based on Particle Evolution
- Be Precise or Fuzzy: Learning the Meaning of Cardinals and Quantifiers from Vision
- Constructing Multiple Tasks for Augmentation: Improving Neural Image Classification With K-means Features
- Track Facial Points in Unconstrained Videos
- Discovering Visual Concept Structure with Sparse and Incomplete Tags
- Automatic Visual Theme Discovery from Joint Image and Text Corpora
- Dynamic Joint Variational Graph Autoencoders
- Extra Proximal-Gradient Inspired Non-local Network
- Convolutional Neural Network on Semi-Regular Triangulated Meshes and its Application to Brain Image Data
- An Effective Training Method For Deep Convolutional Neural Network
- Object Specific Deep Learning Feature and Its Application to Face Detection
- Iterative Object and Part Transfer for Fine-Grained Recognition
- Driver Distraction Identification with an Ensemble of Convolutional Neural Networks
- Superpixel-based Semantic Segmentation Trained by Statistical Process Control
- Use of First and Third Person Views for Deep Intersection Classification
- Efficient Image Splicing Localization via Contrastive Feature Extraction
- Hamiltonian Monte-Carlo for Orthogonal Matrices
- Scalable Annotation of Fine-Grained Categories Without Experts
- Data-Driven Microstructure Property Relations
- Segmentation Free Object Discovery in Video
- A Useful Motif for Flexible Task Learning in an Embodied Two-Dimensional Visual Environment
- Photofeeler-D3: A Neural Network with Voter Modeling for Dating Photo Impression Prediction
- Deep Built-Structure Counting in Satellite Imagery Using Attention Based Re-Weighting
- Scalable Compression of Deep Neural Networks
- RGB-D Individual Segmentation
- Sex-Prediction from Periocular Images across Multiple Sensors and Spectra
- Learning eating environments through scene clustering
- MomentsNet: a simple learning-free method for binary image recognition
- Hard Negative Mining for Metric Learning Based Zero-Shot Classification
- Layout-Graph Reasoning for Fashion Landmark Detection
- Multi-modal dialog for browsing large visual catalogs using exploration-exploitation paradigm in a joint embedding space
- A real-time hourly ozone prediction system using deep convolutional neural network
- ROSA: Robust Salient Object Detection against Adversarial Attacks
- Feature discriminativity estimation in CNNs for transfer learning
- Analyzing Learned Convnet Features with Dirichlet Process Gaussian Mixture Models
- A new neural-network-based model for measuring the strength of a pseudorandom binary sequence
- Detecting Lesion Bounding Ellipses With Gaussian Proposal Networks
- ROI Pooled Correlation Filters for Visual Tracking
- Learning Dilation Factors for Semantic Segmentation of Street Scenes
- Effect of Various Regularizers on Model Complexities of Neural Networks in Presence of Input Noise
- DropRegion Training of Inception Font Network for High-Performance Chinese Font Recognition
- Classificação de espécies de peixe utilizando redes neurais convolucional
- Learning Cascaded Siamese Networks for High Performance Visual Tracking
- Perception-oriented Single Image Super-Resolution via Dual Relativistic Average Generative Adversarial Networks
- Deep rank-based transposition-invariant distances on musical sequences
- Per-Pixel Feedback for improving Semantic Segmentation
- An End-to-End Approach to Automatic Speech Assessment for Cantonese-speaking People with Aphasia
- Compression Artifacts Reduction by a Deep Convolutional Network
- Modular Continual Learning in a Unified Visual Environment
- Fast Predictive Image Registration
- Unsupervised Triplet Hashing for Fast Image Retrieval
- Adaptive Gradient for Adversarial Perturbations Generation
- Saliency Detection via Combining Region-Level and Pixel-Level Predictions with CNNs
- Gaussian Filter in CRF Based Semantic Segmentation
- Understanding the Importance of Single Directions via Representative Substitution
- MIML-FCN+: Multi-instance Multi-label Learning via Fully Convolutional Networks with Privileged Information
- In-Place Zero-Space Memory Protection for CNN
- Dynamic Regularizer with an Informative Prior
- A Shallow High-Order Parametric Approach to Data Visualization and Compression
- On Class Imbalance and Background Filtering in Visual Relationship Detection
- ON-TRAC Consortium End-to-End Speech Translation Systems for the IWSLT 2019 Shared Task
- Hierarchical Scene Parsing by Weakly Supervised Learning with Image Descriptions
- Pyramidal RoR for Image Classification
- Generating Adversarial Examples With Conditional Generative Adversarial Net
- MICIK: MIning Cross-Layer Inherent Similarity Knowledge for Deep Model Compression
- Deep Diagnostics: Applying Convolutional Neural Networks for Vessels Defects Detection
- Literature Review: Human Segmentation with Static Camera
- Full-stack Optimization for Accelerating CNNs with FPGA Validation
- Bi-stream Pose Guided Region Ensemble Network for Fingertip Localization from Stereo Images
- CNNs are Globally Optimal Given Multi-Layer Support
- CNN-based Semantic Segmentation using Level Set Loss
- Multimodal Image Outpainting With Regularized Normalized Diversification
- Synthesizing New Retinal Symptom Images by Multiple Generative Models
- Using Deep Cross Modal Hashing and Error Correcting Codes for Improving the Efficiency of Attribute Guided Facial Image Retrieval
- Circulant Binary Convolutional Networks: Enhancing the Performance of 1-bit DCNNs with Circulant Back Propagation
- ProLFA: Representative Prototype Selection for Local Feature Aggregation
- Assisting human experts in the interpretation of their visual process: A case study on assessing copper surface adhesive potency
- Relative Depth Order Estimation Using Multi-scale Densely Connected Convolutional Networks
- Learning Joint Representations of Videos and Sentences with Web Image Search
- FHEDN: A based on context modeling Feature Hierarchy Encoder-Decoder Network for face detection
- Machine Learning on Biomedical Images: Interactive Learning, Transfer Learning, Class Imbalance, and Beyond
- Improving drone localisation around wind turbines using monocular model-based tracking
- Learning to see across Domains and Modalities
- Can We Automate Diagrammatic Reasoning?
- Discriminate-and-Rectify Encoders: Learning from Image Transformation Sets
- SafeNet: An Assistive Solution to Assess Incoming Threats for Premises
- Fusing Deep Convolutional Networks for Large Scale Visual Concept Classification
- CS591 Report: Application of siamesa network in 2D transformation
- Visualizing and Describing Fine-grained Categories as Textures
- Scene Text Recognition With Finer Grid Rectification
- Core Sampling Framework for Pixel Classification
- Predictive Coding Networks Meet Action Recognition
- Deep 3D Pan via adaptive "t-shaped" convolutions with global and local adaptive dilations
- Dynamic Error-bounded Lossy Compression (EBLC) to Reduce the Bandwidth Requirement for Real-time Vision-based Pedestrian Safety Applications
- Assessing Post Deletion in Sina Weibo: Multi-modal Classification of Hot Topics
- Boosting Mapping Functionality of Neural Networks via Latent Feature Generation based on Reversible Learning
- Zero-shot Learning of 3D Point Cloud Objects
- Purifying Naturalistic Images through a Real-time Style Transfer Semantics Network
- Self-Adaptive Network Pruning
- Towards Interpretable Deep Extreme Multi-label Learning
- Adversarial Incremental Learning
- Analyzing the Dependency of ConvNets on Spatial Information
- Geometry-Based Region Proposals for Real-Time Robot Detection of Tabletop Objects
- Appearance-Based Gaze Estimation Using Dilated-Convolutions
- Brno Mobile OCR Dataset
- Forensic Scanner Identification Using Machine Learning
- Effects of annotation granularity in deep learning models for histopathological images
- ILCRO: Making Importance Landscapes Flat Again
- Re-learning of Child Model for Misclassified data by using KL Divergence in AffectNet: A Database for Facial Expression
- Multiple VLAD encoding of CNNs for image classification
- Non-destructive three-dimensional measurement of hand vein based on self-supervised network
- Multi-View Surveillance Video Summarization via Joint Embedding and Sparse Optimization
- A Deep Optimization Approach for Image Deconvolution
- A Survey On 3D Inner Structure Prediction from its Outer Shape
- Detecting Deep Neural Network Defects with Data Flow Analysis
- Deep Convolutional Poses for Human Interaction Recognition in Monocular Videos
- BOBBY2: Buffer Based Robust High-Speed Object Tracking
- Object Detection on Single Monocular Images through Canonical Correlation Analysis
- Acoustic Scene Classification Using Bilinear Pooling on Time-liked and Frequency-liked Convolution Neural Network
- Hierarchical Feature-Aware Tracking
- Improving Catheter Segmentation & Localization in 3D Cardiac Ultrasound Using Direction-Fused FCN
- Counting dense objects in remote sensing images
- Patch redundancy in images: a statistical testing framework and some applications
- An Internal Covariate Shift Bounding Algorithm for Deep Neural Networks by Unitizing Layers' Outputs
- Phenotypic Profiling of High Throughput Imaging Screens with Generic Deep Convolutional Features
- An End-to-End Framework for Unsupervised Pose Estimation of Occluded Pedestrians
- Residual Continual Learning
- Did Evolution get it right? An evaluation of Near-Infrared imaging in semantic scene segmentation using deep learning
- An Improved Neural Segmentation Method Based on U-NET
- 'Part'ly first among equals: Semantic part-based benchmarking for state-of-the-art object recognition systems
- Learning to Inpaint by Progressively Growing the Mask Regions
- A Targeted Acceleration and Compression Framework for Low bit Neural Networks
- Diversity Transfer Network for Few-Shot Learning
- CHD:Consecutive Horizontal Dropout for Human Gait Feature Extraction
- Using colorization as a tool for automatic makeup suggestion
- PoET-BiN: Power Efficient Tiny Binary Neurons
- ADA-Tucker: Compressing Deep Neural Networks via Adaptive Dimension Adjustment Tucker Decomposition
- Similarities and differences between stimulus tuning in the inferotemporal visual cortex and convolutional networks
- Estimating the Operating Characteristics of Ensemble Methods
- Pipelined Training with Stale Weights of Deep Convolutional Neural Networks
- Network-Density-Controlled Decentralized Parallel Stochastic Gradient Descent in Wireless Systems
- ParNet: Position-aware Aggregated Relation Network for Image-Text matching
- Improving Multi-label Learning with Missing Labels by Structured Semantic Correlations
- Multi-Scale Convolutions for Learning Context Aware Feature Representations
- Mixture separability loss in a deep convolutional network for image classification
- Defending Against Adversarial Attacks Using Random Forests
- Residual Switching Network for Portfolio Optimization
- MV-C3D: A Spatial Correlated Multi-View 3D Convolutional Neural Networks
- Modelling response to trypophobia trigger using intermediate layers of ImageNet networks
- Joint 2D-3D Breast Cancer Classification
- Learning a Complete Image Indexing Pipeline
- On evaluating CNN representations for low resource medical image classification
- Black Box Algorithm Selection by Convolutional Neural Network
- Feature Fusion Detector for Semantic Cognition of Remote Sensing
- Enhanced generative adversarial network for 3D brain MRI super-resolution
- An EEG-based Image Annotation System
- DeepIlluminance: Contextual Illuminance Estimation via Deep Neural Networks
- Improve SGD Training via Aligning Mini-batches
- A deep learning framework for morphologic detail beyond the diffraction limit in infrared spectroscopic imaging
- Semantic Example Guided Image-to-Image Translation
- Semi-Automatic Crowdsourcing Tool for Online Food Image Collection and Annotation
- Capturing Localized Image Artifacts through a CNN-based Hyper-image Representation
- Research Frontiers in Transfer Learning -- a systematic and bibliometric review
- Identity Recognition in Intelligent Cars with Behavioral Data and LSTM-ResNet Classifier
- Weeping and Gnashing of Teeth: Teaching Deep Learning in Image and Video Processing Classes
- Quality-aware Unpaired Image-to-Image Translation
- Network Horizon Dynamics I: Qualitative Aspects
- Region-based semantic segmentation with end-to-end training
- What Else Can Fool Deep Learning? Addressing Color Constancy Errors on Deep Neural Network Performance
- Shape Constrained Network for Eye Segmentation in the Wild
- Simultaneously Learning Architectures and Features of Deep Neural Networks
- Interpretable Deep Neural Networks for Facial Expression and Dimensional Emotion Recognition in-the-wild
- Predicting city safety perception based on visual image content
- TAN: Temporal Aggregation Network for Dense Multi-label Action Recognition
- Graph Interpolating Activation Improves Both Natural and Robust Accuracies in Data-Efficient Deep Learning
- Unconstrained Road Marking Recognition with Generative Adversarial Networks
- Weakly supervised segment annotation via expectation kernel density estimation
- Preliminary study on the modal decomposition of Hermite Gaussian beams via deep learning
- Large-scale Kernel Methods and Applications to Lifelong Robot Learning