Caffe: Convolutional Architecture for Fast Feature Embedding
arXiv:1408.5093
Abstract
Caffe provides multimedia scientists and practitioners with a clean and modifiable framework for state-of-the-art deep learning algorithms and a collection of reference models. The framework is a BSD-licensed C++ library with Python and MATLAB bindings for training and deploying general-purpose convolutional neural networks and other deep models efficiently on commodity architectures. Caffe fits industry and internet-scale media needs by CUDA GPU computation, processing over 40 million images a day on a single K40 or Titan GPU ( 2.5 ms per image). By separating model representation from actual implementation, Caffe allows experimentation and seamless switching among platforms for ease of development and deployment from prototyping machines to cloud environments. Caffe is maintained and developed by the Berkeley Vision and Learning Center (BVLC) with the help of an active community of contributors on GitHub. It powers ongoing research projects, large-scale industrial applications, and startup prototypes in vision, speech, and multimedia.
Tech report for the Caffe software at http://github.com/BVLC/Caffe/
Cited by in corpus (462)
- Convolutional Neural Networks for Medical Image Analysis: Full Training or Fine Tuning?
- Unsupervised Domain Adaptation by Backpropagation
- Striving for Simplicity: The All Convolutional Net
- Understanding Neural Networks Through Deep Visualization
- Explainable Artificial Intelligence: Understanding, Visualizing and Interpreting Deep Learning Models
- cuDNN: Efficient Primitives for Deep Learning
- ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices
- Trained Ternary Quantization
- Network Trimming: A Data-Driven Neuron Pruning Approach towards Efficient Deep Architectures
- FCNs in the Wild: Pixel-level Adversarial and Constraint-based Adaptation
- Automatic Liver and Lesion Segmentation in CT Using Cascaded Fully Convolutional Neural Networks and 3D Conditional Random Fields
- Learning Deconvolution Network for Semantic Segmentation
- FlowNet: Learning Optical Flow with Convolutional Networks
- Focusing Attention: Towards Accurate Text Recognition in Natural Images
- Training Convolutional Networks with Noisy Labels
- Boosting the accuracy of multi-spectral image pan-sharpening by learning a deep residual network
- Learning Structured Sparsity in Deep Neural Networks
- Adversarial Discriminative Domain Adaptation
- Visualizing Deep Neural Network Decisions: Prediction Difference Analysis
- GLAD: Global-Local-Alignment Descriptor for Pedestrian Retrieval
- Cascade R-CNN: Delving into High Quality Object Detection
- Deep CORAL: Correlation Alignment for Deep Domain Adaptation
- DeepNAT: Deep Convolutional Neural Network for Segmenting Neuroanatomy
- Channel Pruning for Accelerating Very Deep Neural Networks
- Learning Activation Functions to Improve Deep Neural Networks
- A Convolutional Neural Network Approach for Post-Processing in HEVC Intra Coding
- Speeding-up Convolutional Neural Networks Using Fine-tuned CP-Decomposition
- Visual Saliency Based on Multiscale Deep Features
- Fast Convolutional Nets With fbfft: A GPU Performance Evaluation
- Detecting Sarcasm in Multimodal Social Platforms
- Deep Gaze I: Boosting Saliency Prediction with Feature Maps Trained on ImageNet
- Exploring the Regularity of Sparse Structure in Convolutional Neural Networks
- Superhuman Accuracy on the SNEMI3D Connectomics Challenge
- Scene Text Detection via Holistic, Multi-Channel Prediction
- MemNet: A Persistent Memory Network for Image Restoration
- Fathom: Reference Workloads for Modern Deep Learning Methods
- Pose Invariant Embedding for Deep Person Re-identification
- Learning Visual Importance for Graphic Designs and Data Visualizations
- FusionNet: 3D Object Classification Using Multiple Data Representations
- Training Deeper Convolutional Networks with Deep Supervision
- Breast Mass Classification from Mammograms using Deep Convolutional Neural Networks
- Deeply-Learned Part-Aligned Representations for Person Re-Identification
- Exploring Nearest Neighbor Approaches for Image Captioning
- End-to-End Photo-Sketch Generation via Fully Convolutional Representation Learning
- Dynamic Network Surgery for Efficient DNNs
- An All-in-One Network for Dehazing and Beyond
- ImageNet pre-trained models with batch normalization
- A Pursuit of Temporal Accuracy in General Activity Detection
- Memory Bounded Deep Convolutional Networks
- Deep Transfer Network: Unsupervised Domain Adaptation
- CrowdNet: A Deep Convolutional Network for Dense Crowd Counting
- DeepSketch2Face: A Deep Learning Based Sketching System for 3D Face and Caricature Modeling
- Amulet: Aggregating Multi-level Convolutional Features for Salient Object Detection
- VoxResNet: Deep Voxelwise Residual Networks for Volumetric Brain Segmentation
- Single-Shot Refinement Neural Network for Object Detection
- Learning Deep Structured Multi-Scale Features using Attention-Gated CRFs for Contour Prediction
- Multi-modal Factorized Bilinear Pooling with Co-Attention Learning for Visual Question Answering
- Ultimate tensorization: compressing convolutional and FC layers alike
- Wavelet Convolutional Neural Networks for Texture Classification
- Learning a visuomotor controller for real world robotic grasping using simulated depth images
- Direct Multitype Cardiac Indices Estimation via Joint Representation and Regression Learning
- PixelNet: Representation of the pixels, by the pixels, and for the pixels
- A Unified Perspective on Multi-Domain and Multi-Task Learning
- Universal adversarial perturbations
- Learning Deep Context-aware Features over Body and Latent Parts for Person Re-identification
- Perceptual Generative Adversarial Networks for Small Object Detection
- Memory-Efficient Implementation of DenseNets
- Deep Convolutional Neural Network Inference with Floating-point Weights and Fixed-point Activations
- Exploiting Local Features from Deep Networks for Image Retrieval
- Deep Probabilistic Programming
- Model-based Deep Hand Pose Estimation
- Stacked Deconvolutional Network for Semantic Segmentation
- Clipper: A Low-Latency Online Prediction Serving System
- Deep Neural Networks are Easily Fooled: High Confidence Predictions for Unrecognizable Images
- A Computer Vision System to Localize and Classify Wastes on the Streets
- Image Prediction for Limited-angle Tomography via Deep Learning with Convolutional Neural Network
- NeuralPower: Predict and Deploy Energy-Efficient Convolutional Neural Networks
- Towards Cognitive Exploration through Deep Reinforcement Learning for Mobile Robots
- Hierarchical Attention Network for Action Recognition in Videos
- Towards seamless multi-view scene analysis from satellite to street-level
- Deep Convolution Networks for Compression Artifacts Reduction
- Learning Language-Visual Embedding for Movie Understanding with Natural-Language
- Learning Background-Aware Correlation Filters for Visual Tracking
- Benchmarking State-of-the-Art Deep Learning Software Tools
- Illuminating Pedestrians via Simultaneous Detection & Segmentation
- Highly Efficient Forward and Backward Propagation of Convolutional Neural Networks for Pixelwise Classification
- CREST: Convolutional Residual Learning for Visual Tracking
- See, Hear, and Read: Deep Aligned Representations
- Background rejection in NEXT using deep neural networks
- Deep Learning with Low Precision by Half-wave Gaussian Quantization
- TensorLayer: A Versatile Library for Efficient Deep Learning Development
- Single Shot Text Detector with Regional Attention
- Implementing the Deep Q-Network
- Parallel Tracking and Verifying: A Framework for Real-Time and High Accuracy Visual Tracking
- Learning Spatial Regularization with Image-level Supervisions for Multi-label Image Classification
- PixelNet: Towards a General Pixel-level Architecture
- Online Multi-Object Tracking Using CNN-based Single Object Tracker with Spatial-Temporal Attention Mechanism
- Deep Direct Regression for Multi-Oriented Scene Text Detection
- ShiftCNN: Generalized Low-Precision Architecture for Inference of Convolutional Neural Networks
- Multi-label Image Recognition by Recurrently Discovering Attentional Regions
- End-to-End Image Super-Resolution via Deep and Shallow Convolutional Networks
- Deep learning for plasma tomography using the bolometer system at JET
- Automatic 3D Cardiovascular MR Segmentation with Densely-Connected Volumetric ConvNets
- Label Refinement Network for Coarse-to-Fine Semantic Segmentation
- Range Loss for Deep Face Recognition with Long-tail
- The Freiburg Groceries Dataset
- A Localisation-Segmentation Approach for Multi-label Annotation of Lumbar Vertebrae using Deep Nets
- Interleaved Text/Image Deep Mining on a Large-Scale Radiology Database for Automated Image Interpretation
- YouTube-BoundingBoxes: A Large High-Precision Human-Annotated Data Set for Object Detection in Video
- DeepVO: A Deep Learning approach for Monocular Visual Odometry
- ImageNet Training in Minutes
- Deep Temporal Appearance-Geometry Network for Facial Expression Recognition
- SemanticFusion: Dense 3D Semantic Mapping with Convolutional Neural Networks
- Low-memory GEMM-based convolution algorithms for deep neural networks
- Feedforward semantic segmentation with zoom-out features
- Semantic Scene Completion from a Single Depth Image
- Graph Based Convolutional Neural Network
- Incorporating Network Built-in Priors in Weakly-supervised Semantic Segmentation
- A deep architecture for unified aesthetic prediction
- Learning Deep Similarity Models with Focus Ranking for Fabric Image Retrieval
- Shift: A Zero FLOP, Zero Parameter Alternative to Spatial Convolutions
- Photo Aesthetics Ranking Network with Attributes and Content Adaptation
- Boosting Image Captioning with Attributes
- Deep FisherNet for Object Classification
- FaceNet2ExpNet: Regularizing a Deep Face Recognition Net for Expression Recognition
- Pose-driven Deep Convolutional Model for Person Re-identification
- Learning to Inpaint for Image Compression
- Multi-scale Deep Learning Architectures for Person Re-identification
- Weakly- and Semi-Supervised Object Detection with Expectation-Maximization Algorithm
- On the Robustness of Convolutional Neural Networks to Internal Architecture and Weight Perturbations
- Fashion DNA: Merging Content and Sales Data for Recommendation and Article Mapping
- SC-DCNN: Highly-Scalable Deep Convolutional Neural Network using Stochastic Computing
- Understanding and Comparing Deep Neural Networks for Age and Gender Classification
- Deep Neural Network with l2-norm Unit for Brain Lesions Detection
- SFD: Single Shot Scale-invariant Face Detector
- Deep Joint Face Hallucination and Recognition
- Discriminative Unsupervised Feature Learning with Exemplar Convolutional Neural Networks
- Multi-Modality Fusion based on Consensus-Voting and 3D Convolution for Isolated Gesture Recognition
- Distributed Robust Learning
- Speeding up Convolutional Neural Networks By Exploiting the Sparsity of Rectifier Units
- Deep Ordinal Ranking for Multi-Category Diagnosis of Alzheimer's Disease using Hippocampal MRI data
- Incorporating Copying Mechanism in Image Captioning for Learning Novel Objects
- Face Parsing via a Fully-Convolutional Continuous CRF Neural Network
- Fisher Kernel for Deep Neural Activations
- Classification regions of deep neural networks
- The Cross-Depiction Problem: Computer Vision Algorithms for Recognising Objects in Artwork and in Photographs
- On Vectorization of Deep Convolutional Neural Networks for Vision Tasks
- Adapting Deep Network Features to Capture Psychological Representations
- Visual Servoing from Deep Neural Networks
- Towards Automatic Abdominal Multi-Organ Segmentation in Dual Energy CT using Cascaded 3D Fully Convolutional Network
- End-to-end people detection in crowded scenes
- Automatic Discovery and Optimization of Parts for Image Classification
- Freehand Sketch Recognition Using Deep Features
- Convolutional Neural Networks at Constrained Time Cost
- Joint Object and Part Segmentation using Deep Learned Potentials
- Unsupervised Representation Learning by Sorting Sequences
- Learning to Prune Filters in Convolutional Neural Networks
- Face Detection using Deep Learning: An Improved Faster RCNN Approach
- A Tour of TensorFlow
- Viewpoints and Keypoints
- RankIQA: Learning from Rankings for No-reference Image Quality Assessment
- Learning Affinity via Spatial Propagation Networks
- Sliding Line Point Regression for Shape Robust Scene Text Detection
- Pose-Invariant Face Alignment with a Single CNN
- Learning Accurate Low-Bit Deep Neural Networks with Stochastic Quantization
- Single- and Multi-Task Architectures for Surgical Workflow Challenge at M2CAI 2016
- Interpreting the Predictions of Complex ML Models by Layer-wise Relevance Propagation
- Convergent Block Coordinate Descent for Training Tikhonov Regularized Deep Neural Networks
- PupilNet v2.0: Convolutional Neural Networks for CPU based real time Robust Pupil Detection
- Deep Sketch Hashing: Fast Free-hand Sketch-Based Image Retrieval
- Zoom Out-and-In Network with Recursive Training for Object Proposal
- Single- and Multi-Task Architectures for Tool Presence Detection Challenge at M2CAI 2016
- Matching-CNN Meets KNN: Quasi-Parametric Human Parsing
- Learning to count with deep object features
- CASENet: Deep Category-Aware Semantic Edge Detection
- Factors of Transferability for a Generic ConvNet Representation
- Detecting Vanishing Points using Global Image Context in a Non-Manhattan World
- Caffe con Troll: Shallow Ideas to Speed Up Deep Learning
- Model compression as constrained optimization, with application to neural nets. Part II: quantization
- Attention-Based Multimodal Fusion for Video Description
- Recovering the Missing Link: Predicting Class-Attribute Associations for Unsupervised Zero-Shot Learning
- Flower Categorization using Deep Convolutional Neural Networks
- Attentional Network for Visual Object Detection
- Localizing and Orienting Street Views Using Overhead Imagery
- Not All Ops Are Created Equal!
- Material Recognition in the Wild with the Materials in Context Database
- Neural Person Search Machines
- Cognitive Database: A Step towards Endowing Relational Databases with Artificial Intelligence Capabilities
- Learning Deep Networks from Noisy Labels with Dropout Regularization
- Deep Image Harmonization
- Evaluating Content-centric vs User-centric Ad Affect Recognition
- Depth2Action: Exploring Embedded Depth for Large-Scale Action Recognition
- Are Safer Looking Neighborhoods More Lively? A Multimodal Investigation into Urban Life
- Can we unify monocular detectors for autonomous driving by using the pixel-wise semantic segmentation of CNNs?
- Iterative Multi-domain Regularized Deep Learning for Anatomical Structure Detection and Segmentation from Ultrasound Images
- Improving utility of brain tumor confocal laser endomicroscopy: objective value assessment and diagnostic frame detection with convolutional neural networks
- Image Credibility Analysis with Effective Domain Transferred Deep Networks
- Self-supervised learning of visual features through embedding images into text topic spaces
- Deep Structured Models For Group Activity Recognition
- Deep Cuboid Detection: Beyond 2D Bounding Boxes
- HashGAN:Attention-aware Deep Adversarial Hashing for Cross Modal Retrieval
- Modeling Spatial and Temporal Cues for Multi-label Facial Action Unit Detection
- Half-CNN: A General Framework for Whole-Image Regression
- Where to Focus: Query Adaptive Matching for Instance Retrieval Using Convolutional Feature Maps
- Treelogy: A Novel Tree Classifier Utilizing Deep and Hand-crafted Representations
- Tartan: Accelerating Fully-Connected and Convolutional Layers in Deep Learning Networks by Exploiting Numerical Precision Variability
- Deep Predictive Policy Training using Reinforcement Learning
- Learning Hypergraph-regularized Attribute Predictors
- Learning Multi-Scale Deep Features for High-Resolution Satellite Image Classification
- Privacy-Preserving Deep Inference for Rich User Data on The Cloud
- Scene Parsing with Global Context Embedding
- Why my photos look sideways or upside down? Detecting Canonical Orientation of Images using Convolutional Neural Networks
- A New Convolutional Network-in-Network Structure and Its Applications in Skin Detection, Semantic Segmentation, and Artifact Reduction
- Deep Embedding Convolutional Neural Network for Synthesizing CT Image from T1-Weighted MR Image
- Direct Estimation of Regional Wall Thicknesses via Residual Recurrent Neural Network
- Ligand Pose Optimization with Atomic Grid-Based Convolutional Neural Networks
- Contextual Multi-Scale Region Convolutional 3D Network for Activity Detection
- Comparing Neural and Attractiveness-based Visual Features for Artwork Recommendation
- Let's Dance: Learning From Online Dance Videos
- Exploring Food Detection using CNNs
- Transfer Learning for Material Classification using Convolutional Networks
- Contrastive-center loss for deep neural networks
- Convolutional Neural Network-Based Image Representation for Visual Loop Closure Detection
- Inverse Compositional Spatial Transformer Networks
- Exploring the Design Space of Deep Convolutional Neural Networks at Large Scale
- Cardea: Context-Aware Visual Privacy Protection from Pervasive Cameras
- Action Recognition with Joint Attention on Multi-Level Deep Features
- Visual Discovery at Pinterest
- Feature Incay for Representation Regularization
- Joint Prediction of Depths, Normals and Surface Curvature from RGB Images using CNNs
- Active Convolution: Learning the Shape of Convolution for Image Classification
- DeepEdge: A Multi-Scale Bifurcated Deep Network for Top-Down Contour Detection
- Multi-task Dictionary Learning based Convolutional Neural Network for Computer aided Diagnosis with Longitudinal Images
- Learning Common and Specific Features for RGB-D Semantic Segmentation with Deconvolutional Networks
- Deep Heterogeneous Feature Fusion for Template-Based Face Recognition
- Enabling My Robot To Play Pictionary : Recurrent Neural Networks For Sketch Recognition
- Deep Temporal Linear Encoding Networks
- Object Level Deep Feature Pooling for Compact Image Representation
- Situational Object Boundary Detection
- Weakly-supervised Learning of Mid-level Features for Pedestrian Attribute Recognition and Localization
- BT-Nets: Simplifying Deep Neural Networks via Block Term Decomposition
- An Error Detection and Correction Framework for Connectomics
- From Plants to Landmarks: Time-invariant Plant Localization that uses Deep Pose Regression in Agricultural Fields
- Deep Learning for identifying radiogenomic associations in breast cancer
- Actions and Attributes from Wholes and Parts
- Learning RGB-D Salient Object Detection using background enclosure, depth contrast, and top-down features
- Reconstructing Vechicles from a Single Image: Shape Priors for Road Scene Understanding
- Deep Joint Rain Detection and Removal from a Single Image
- A Jointly Learned Deep Architecture for Facial Attribute Analysis and Face Detection in the Wild
- A New Evaluation Protocol and Benchmarking Results for Extendable Cross-media Retrieval
- Multi-Object Classification and Unsupervised Scene Understanding Using Deep Learning Features and Latent Tree Probabilistic Models
- Estimated Depth Map Helps Image Classification
- Learning a Dilated Residual Network for SAR Image Despeckling
- Zoom-in-Net: Deep Mining Lesions for Diabetic Retinopathy Detection
- Content-Based Video Retrieval in Historical Collections of the German Broadcasting Archive
- BMXNet: An Open-Source Binary Neural Network Implementation Based on MXNet
- Cultural Event Recognition with Visual ConvNets and Temporal Models
- Collaborative Summarization of Topic-Related Videos
- Learning with Hierarchical Gaussian Kernels
- Learning a Repression Network for Precise Vehicle Search
- FoveaNet: Perspective-aware Urban Scene Parsing
- MPIIGaze: Real-World Dataset and Deep Appearance-Based Gaze Estimation
- Dual Path Networks for Multi-Person Human Pose Estimation
- 3D Reconstruction in Canonical Co-ordinate Space from Arbitrarily Oriented 2D Images
- Computer-Aided Colorectal Tumor Classification in NBI Endoscopy Using CNN Features
- Comprehensive Evaluation of OpenCL-based Convolutional Neural Network Accelerators in Xilinx and Altera FPGAs
- Adaptive Feeding: Achieving Fast and Accurate Detections by Adaptively Combining Object Detectors
- Learning Robust Deep Face Representation
- Distributed Training Large-Scale Deep Architectures
- What Is the Best Practice for CNNs Applied to Visual Instance Retrieval?
- A 4D Light-Field Dataset and CNN Architectures for Material Recognition
- Fast and Energy-Efficient CNN Inference on IoT Devices
- Weaving Multi-scale Context for Single Shot Detector
- Leaf Identification Using a Deep Convolutional Neural Network
- Social Behavior Prediction from First Person Videos
- Fine-grained Recognition in the Wild: A Multi-Task Domain Adaptation Approach
- Optimized Broadcast for Deep Learning Workloads on Dense-GPU InfiniBand Clusters: MPI or NCCL?
- Two-stream Flow-guided Convolutional Attention Networks for Action Recognition
- All-Transfer Learning for Deep Neural Networks and its Application to Sepsis Classification
- Performance Evaluation of Deep Learning Tools in Docker Containers
- Texture segmentation with Fully Convolutional Networks
- Pose-Aware Person Recognition
- nuts-flow/ml: data pre-processing for deep learning
- Optimizing Region Selection for Weakly Supervised Object Detection
- Spatio-temporal Human Action Localisation and Instance Segmentation in Temporally Untrimmed Videos
- Simultaneous Hand Pose and Skeleton Bone-Lengths Estimation from a Single Depth Image
- FOTS: Fast Oriented Text Spotting with a Unified Network
- Surveillance Video Parsing with Single Frame Supervision
- Photo Filter Recommendation by Category-Aware Aesthetic Learning
- Understanding the Impact of Precision Quantization on the Accuracy and Energy of Neural Networks
- Object-Scene Convolutional Neural Networks for Event Recognition in Images
- Deep 3D Face Identification
- Outlier Robust Online Learning
- Convolutional Regression for Visual Tracking
- Deep Scene Text Detection with Connected Component Proposals
- Becoming the Expert - Interactive Multi-Class Machine Teaching
- Real-time Human Pose Estimation from Video with Convolutional Neural Networks
- Deep Quantization: Encoding Convolutional Activations with Deep Generative Model
- Features in Concert: Discriminative Feature Selection meets Unsupervised Clustering
- Fast Deep Matting for Portrait Animation on Mobile Phone
- Deep Feature Learning via Structured Graph Laplacian Embedding for Person Re-Identification
- Spatial Pyramid Convolutional Neural Network for Social Event Detection in Static Image
- Context-Aware Semantic Inpainting
- How hard can it be? Estimating the difficulty of visual search in an image
- DeepSkeleton: Skeleton Map for 3D Human Pose Regression
- Multi-stage Object Detection with Group Recursive Learning
- Modeling Multimodal Clues in a Hybrid Deep Learning Framework for Video Classification
- Tracking Persons-of-Interest via Unsupervised Representation Adaptation
- HMD Vision-based Teleoperating UGV and UAV for Hostile Environment using Deep Learning
- Computation Error Analysis of Block Floating Point Arithmetic Oriented Convolution Neural Network Accelerator Design
- DeepFace: Face Generation using Deep Learning
- Visual Summary of Egocentric Photostreams by Representative Keyframes
- Predicting Aesthetic Score Distribution through Cumulative Jensen-Shannon Divergence
- Action Classification and Highlighting in Videos
- Deep Convolutional Neural Network for 6-DOF Image Localization
- Multi-Task Curriculum Transfer Deep Learning of Clothing Attributes
- Visual Stability Prediction and Its Application to Manipulation
- VIPLFaceNet: An Open Source Deep Face Recognition SDK
- Searching Action Proposals via Spatial Actionness Estimation and Temporal Path Inference and Tracking
- Commonly Uncommon: Semantic Sparsity in Situation Recognition
- Flexible Deep Neural Network Processing
- Bridging the Gap Between Neural Networks and Neuromorphic Hardware with A Neural Network Compiler
- Large-scale Isolated Gesture Recognition Using Convolutional Neural Networks
- SparCE: Sparsity aware General Purpose Core Extensions to Accelerate Deep Neural Networks
- Learning Gating ConvNet for Two-Stream based Methods in Action Recognition
- Tell and Predict: Kernel Classifier Prediction for Unseen Visual Classes from Unstructured Text Descriptions
- Group-wise Deep Co-saliency Detection
- Permutohedral Lattice CNNs
- Contour Detection Using Cost-Sensitive Convolutional Neural Networks
- Co-Regularized Deep Representations for Video Summarization
- Modelling Local Deep Convolutional Neural Network Features to Improve Fine-Grained Image Classification
- Predicting Human Interaction via Relative Attention Model
- Image Classification and Retrieval from User-Supplied Tags
- Relaxed Earth Mover's Distances for Chain- and Tree-connected Spaces and their use as a Loss Function in Deep Learning
- Human Pose Estimation from Depth Images via Inference Embedded Multi-task Learning
- Action Recognition with Coarse-to-Fine Deep Feature Integration and Asynchronous Fusion
- Active learning with version spaces for object detection
- Optimizing Filter Size in Convolutional Neural Networks for Facial Action Unit Recognition
- Brain Abnormality Detection by Deep Convolutional Neural Network
- Deep Hybrid Similarity Learning for Person Re-identification
- PageNet: Page Boundary Extraction in Historical Handwritten Documents
- Mixed context networks for semantic segmentation
- Complex Event Recognition from Images with Few Training Examples
- Araguaia Medical Vision Lab at ISIC 2017 Skin Lesion Classification Challenge
- Learning to Disambiguate by Asking Discriminative Questions
- Generalized Deep Image to Image Regression
- A Multi-view RGB-D Approach for Human Pose Estimation in Operating Rooms
- FFT-Based Deep Learning Deployment in Embedded Systems
- SwiDeN : Convolutional Neural Networks For Depiction Invariant Object Recognition
- Material Classification in the Wild: Do Synthesized Training Data Generalise Better than Real-World Training Data?
- Deep Regression Forests for Age Estimation
- Purine: A bi-graph based deep learning framework
- Viewpoint Invariant Action Recognition using RGB-D Videos
- Deep Kinematic Pose Regression
- Zero-Shot Fine-Grained Classification by Deep Feature Learning with Semantics
- Deep joint rain and haze removal from single images
- Pooled Motion Features for First-Person Videos
- Deep learning analysis of breast MRIs for prediction of occult invasive disease in ductal carcinoma in situ
- Optimal deep neural networks for sparse recovery via Laplace techniques
- Less-forgetful Learning for Domain Expansion in Deep Neural Networks
- Effective Image Retrieval via Multilinear Multi-index Fusion
- Real Time Fine-Grained Categorization with Accuracy and Interpretability
- Multi-Branch Fully Convolutional Network for Face Detection
- A Metaprogramming and Autotuning Framework for Deploying Deep Learning Applications
- Efficient Image Evidence Analysis of CNN Classification Results
- Learning Temporal Embeddings for Complex Video Analysis
- Ensemble of Part Detectors for Simultaneous Classification and Localization
- Personality, Culture, and System Factors - Impact on Affective Response to Multimedia
- Cloudroid: A Cloud Framework for Transparent and QoS-aware Robotic Computation Outsourcing
- Revisiting IM2GPS in the Deep Learning Era
- Design of a Very Compact CNN Classifier for Online Handwritten Chinese Character Recognition Using DropWeight and Global Pooling
- Deep Optical Flow Estimation Via Multi-Scale Correspondence Structure Learning
- Approximate Policy Iteration for Budgeted Semantic Video Segmentation
- Efficient Inferencing of Compressed Deep Neural Networks
- Recycle deep features for better object detection
- Semantic-level Decentralized Multi-Robot Decision-Making using Probabilistic Macro-Observations
- Recognizing Dynamic Scenes with Deep Dual Descriptor based on Key Frames and Key Segments
- Tutorial on Answering Questions about Images with Deep Learning
- Convolutional Neural Network on Three Orthogonal Planes for Dynamic Texture Classification
- Automatic Handgun Detection Alarm in Videos Using Deep Learning
- Person Re-Identification via Recurrent Feature Aggregation
- Distance to Center of Mass Encoding for Instance Segmentation
- BENCHIP: Benchmarking Intelligence Processors
- MLitB: Machine Learning in the Browser
- Neural Signatures for Licence Plate Re-identification
- Discriminatively Learned Hierarchical Rank Pooling Networks
- Bringing Impressionism to Life with Neural Style Transfer in Come Swim
- Learning Local Shape Descriptors from Part Correspondences With Multi-view Convolutional Networks
- CP-decomposition with Tensor Power Method for Convolutional Neural Networks Compression
- When 3D-Aided 2D Face Recognition Meets Deep Learning: An extended UR2D for Pose-Invariant Face Recognition
- DeMeshNet: Blind Face Inpainting for Deep MeshFace Verification
- Image-based Recommendations on Styles and Substitutes
- Cross-Media Similarity Evaluation for Web Image Retrieval in the Wild
- Persistent Evidence of Local Image Properties in Generic ConvNets
- GM-Net: Learning Features with More Efficiency
- StuffNet: Using 'Stuff' to Improve Object Detection
- Progressive Representation Adaptation for Weakly Supervised Object Localization
- What does fault tolerant Deep Learning need from MPI?
- weedNet: Dense Semantic Weed Classification Using Multispectral Images and MAV for Smart Farming
- Segmentation-by-Detection: A Cascade Network for Volumetric Medical Image Segmentation
- Stochastic Gradient MCMC with Stale Gradients
- A generic and fast C++ optimization framework
- Wearing Many (Social) Hats: How Different are Your Different Social Network Personae?
- High Efficient Reconstruction of Single-shot T2 Mapping from OverLapping-Echo Detachment Planar Imaging Based on Deep Residual Network
- Multi-Path Region-Based Convolutional Neural Network for Accurate Detection of Unconstrained "Hard Faces"
- Multi-scale recognition with DAG-CNNs
- Cavs: A Vertex-centric Programming Interface for Dynamic Neural Networks
- Multi-Mode Inference Engine for Convolutional Neural Networks
- Classification Accuracy Improvement for Neuromorphic Computing Systems with One-level Precision Synapses
- Generic 3D Convolutional Fusion for image restoration
- Recognizing and Curating Photo Albums via Event-Specific Image Importance
- Cappuccino: Efficient Inference Software Synthesis for Mobile System-on-Chips
- One-Shot Fine-Grained Instance Retrieval
- Deep, Dense, and Low-Rank Gaussian Conditional Random Fields
- Stacked Approximated Regression Machine: A Simple Deep Learning Approach
- 3D Human Pose Estimation Using Convolutional Neural Networks with 2D Pose Information
- Integrated Deep and Shallow Networks for Salient Object Detection
- Scalable Discrete Supervised Hash Learning with Asymmetric Matrix Factorization
- Chained Multi-stream Networks Exploiting Pose, Motion, and Appearance for Action Classification and Detection
- Progressively Diffused Networks for Semantic Image Segmentation
- IBM Deep Learning Service
- Image Matching via Loopy RNN
- The Compressed Model of Residual CNDS
- Two-stream convolutional neural network for accurate RGB-D fingertip detection using depth and edge information
- Should I use TensorFlow
- End-to-End Data Visualization by Metric Learning and Coordinate Transformation
- Active Learning for Structured Prediction from Partially Labeled Data
- Hardware-Software Codesign of Accurate, Multiplier-free Deep Neural Networks
- Learning and Fusing Multimodal Features from and for Multi-task Facial Computing
- Object Specific Deep Learning Feature and Its Application to Face Detection
- Using Cross-Model EgoSupervision to Learn Cooperative Basketball Intention
- Can Image Retrieval help Visual Saliency Detection?
- Constraint-free Natural Image Reconstruction from fMRI Signals Based on Convolutional Neural Network
- Slim-DP: A Light Communication Data Parallelism for DNN
- Subset Feature Learning for Fine-Grained Category Classification
- A Multi-layered Acoustic Tokenizing Deep Neural Network (MAT-DNN) for Unsupervised Discovery of Linguistic Units and Generation of High Quality Features
- Large-Scale Mapping of Human Activity using Geo-Tagged Videos
- Hierarchical Scene Parsing by Weakly Supervised Learning with Image Descriptions
- Cross-domain Image Retrieval with a Dual Attribute-aware Ranking Network
- Superpixel-based Semantic Segmentation Trained by Statistical Process Control
- Joint Learning of Distributed Representations for Images and Texts
- Prune the Convolutional Neural Networks with Sparse Shrink
- Homomorphic Parameter Compression for Distributed Deep Learning Training
- A CNN Cascade for Landmark Guided Semantic Part Segmentation
- Selfie Detection by Synergy-Constraint Based Convolutional Neural Network
- Handwritten digit string recognition by combination of residual network and RNN-CTC
- Scene Flow to Action Map: A New Representation for RGB-D based Action Recognition with Convolutional Neural Networks
- Video Semantic Object Segmentation by Self-Adaptation of DCNN
- Action Recognition Based on Joint Trajectory Maps Using Convolutional Neural Networks
- When Saliency Meets Sentiment: Understanding How Image Content Invokes Emotion and Sentiment
- Seeing into Darkness: Scotopic Visual Recognition
- Pseudo-positive regularization for deep person re-identification
- Automatic Synchronization of Multi-User Photo Galleries
- Fast Landmark Localization with 3D Component Reconstruction and CNN for Cross-Pose Recognition
- On the Modeling of Error Functions as High Dimensional Landscapes for Weight Initialization in Learning Networks
- Effective Combination of Language and Vision Through Model Composition and the R-CCA Method
- Binary Hashing with Semidefinite Relaxation and Augmented Lagrangian
- Deep Convolutional Poses for Human Interaction Recognition in Monocular Videos
- Towards an "In-the-Wild" Emotion Dataset Using a Game-based Framework
- Comparison of the Deep-Learning-Based Automated Segmentation Methods for the Head Sectioned Images of the Virtual Korean Human Project
- Relative Depth Order Estimation Using Multi-scale Densely Connected Convolutional Networks
- SHOE: Supervised Hashing with Output Embeddings