MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems
arXiv:1512.01274
Abstract
MXNet is a multi-language machine learning (ML) library to ease the development of ML algorithms, especially for deep neural networks. Embedded in the host language, it blends declarative symbolic expression with imperative tensor computation. It offers auto differentiation to derive gradients. MXNet is computation and memory efficient and runs on various heterogeneous systems, ranging from mobile devices to distributed GPU clusters. This paper describes both the API design and the system implementation of MXNet, and explains how embedding of both symbolic expression and tensor operation is handled in a unified fashion. Our preliminary experiments reveal promising results on large scale deep neural network applications using multiple GPU machines.
In Neural Information Processing Systems, Workshop on Machine Learning Systems, 2016
References in corpus (2)
Cited by in corpus (499)
- Array Programming with NumPy
- TensorFlow: A system for large-scale machine learning
- DeePMD-kit: A deep learning package for many-body potential energy representation and molecular dynamics
- ResUNet-a: a deep learning framework for semantic segmentation of remotely sensed data
- Convergence of Edge Computing and Deep Learning: A Comprehensive Survey
- Recent Advances and Applications of Deep Learning Methods in Materials Science
- Deep Learning-Based Communication Over the Air
- A Survey on Distributed Machine Learning
- MMDetection: Open MMLab Detection Toolbox and Benchmark
- Deep Graph Library: A Graph-Centric, Highly-Performant Package for Graph Neural Networks
- The ApolloScape Open Dataset for Autonomous Driving and its Application
- GPyTorch: Blackbox Matrix-Matrix Gaussian Process Inference with GPU Acceleration
- Training Deep Nets with Sublinear Memory Cost
- Open Graph Benchmark: Datasets for Machine Learning on Graphs
- Hyper-Parameter Optimization: A Review of Algorithms and Applications
- Asynchronous Federated Optimization
- SciANN: A Keras/Tensorflow wrapper for scientific computations and physics-informed deep learning using artificial neural networks
- ParticleNet: Jet Tagging via Particle Clouds
- Conditional Positional Encodings for Vision Transformers
- Revisiting Batch Normalization For Practical Domain Adaptation
- Deep Learning for Time Series Forecasting: Tutorial and Literature Survey
- FedML: A Research Library and Benchmark for Federated Machine Learning
- DyNet: The Dynamic Neural Network Toolkit
- Like What You Like: Knowledge Distill via Neuron Selectivity Transfer
- A Survey on Edge Computing Systems and Tools
- GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism
- Learning with a Strong Adversary
- Ray: A Distributed Framework for Emerging AI Applications
- Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions
- Wider or Deeper: Revisiting the ResNet Model for Visual Recognition
- On the Reconstruction of Face Images from Deep Face Templates
- MediaPipe: A Framework for Building Perception Pipelines
- Hardware and Software Optimizations for Accelerating Deep Neural Networks: Survey of Current Trends, Challenges, and the Road Ahead
- GluonCV and GluonNLP: Deep Learning in Computer Vision and Natural Language Processing
- TensorLy: Tensor Learning in Python
- Stochastic Activation Pruning for Robust Adversarial Defense
- Learning to Optimize Tensor Programs
- Pervasive AI for IoT applications: A Survey on Resource-efficient Distributed Artificial Intelligence
- Tensor Methods in Computer Vision and Deep Learning
- Deep Transfer Learning: A new deep learning glitch classification method for advanced LIGO
- Glitch Classification and Clustering for LIGO with Deep Transfer Learning
- A Berkeley View of Systems Challenges for AI
- SuperNeurons: Dynamic GPU Memory Management for Training Deep Neural Networks
- Tiny Machine Learning: Progress and Futures
- D: Decentralized Training over Decentralized Data
- Local Model Poisoning Attacks to Byzantine-Robust Federated Learning
- Generalized Byzantine-tolerant SGD
- MLPerf Training Benchmark
- Simulator-free Solution of High-Dimensional Stochastic Elliptic Partial Differential Equations using Deep Neural Networks
- Deep Learning in the Automotive Industry: Applications and Tools
- Identification of heavy, energetic, hadronically decaying particles using machine-learning techniques
- Improving Neural Network Quantization without Retraining using Outlier Channel Splitting
- Database Meets Deep Learning: Challenges and Opportunities
- Generative Poisoning Attack Method Against Neural Networks
- A nonlocal physics-informed deep learning framework using the peridynamic differential operator
- Supporting Very Large Models using Automatic Dataflow Graph Partitioning
- Beyond Data and Model Parallelism for Deep Neural Networks
- BigDL: A Distributed Deep Learning Framework for Big Data
- Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications
- TVM: An Automated End-to-End Optimizing Compiler for Deep Learning
- Self-Attention Networks for Connectionist Temporal Classification in Speech Recognition
- Scientometric Review of Artificial Intelligence for Operations & Maintenance of Wind Turbines: The Past, Present and Future
- SFace: Sigmoid-Constrained Hypersphere Loss for Robust Face Recognition
- Deep Learning in Mobile and Wireless Networking: A Survey
- Simple Baselines for Human Pose Estimation and Tracking
- Scaling Distributed Machine Learning with In-Network Aggregation
- Multi-style Generative Network for Real-time Transfer
- TF-Ranking: Scalable TensorFlow Library for Learning-to-Rank
- Hardware Acceleration of Sparse and Irregular Tensor Computations of ML Models: A Survey and Insights
- Review: Deep Learning in Electron Microscopy
- Siamese Attentional Keypoint Network for High Performance Visual Tracking
- Torchreid: A Library for Deep Learning Person Re-Identification in Pytorch
- Learning Reinforced Attentional Representation for End-to-End Visual Tracking
- Parameter Hub: a Rack-Scale Parameter Server for Distributed Deep Neural Network Training
- SAR Target Recognition Using the Multi-aspect-aware Bidirectional LSTM Recurrent Neural Networks
- Understanding the effect of hyperparameter optimization on machine learning models for structure design problems
- Demystifying Neural Style Transfer
- Scale-Aware Trident Networks for Object Detection
- Memory-Efficient Implementation of DenseNets
- Neural Models for Information Retrieval
- Evolving Deep Convolutional Neural Networks for Image Classification
- Ansor: Generating High-Performance Tensor Programs for Deep Learning
- Deep Polynomial Neural Networks
- Clipper: A Low-Latency Online Prediction Serving System
- A deep learning framework for solution and discovery in solid mechanics
- Interleaved Group Convolutions for Deep Neural Networks
- DocTer: Documentation Guided Fuzzing for Testing Deep Learning API Functions
- Tensor Regression Networks
- GossipGraD: Scalable Deep Learning using Gossip Communication based Asynchronous Gradient Descent
- Bayesian Layers: A Module for Neural Network Uncertainty
- GluonTS: Probabilistic Time Series Models in Python
- RETURNN as a Generic Flexible Neural Toolkit with Application to Translation and Speech Recognition
- Secure Evaluation of Quantized Neural Networks
- Optimizing CNN Model Inference on CPUs
- LEEP: A New Measure to Evaluate Transferability of Learned Representations
- Interaction networks for the identification of boosted decays
- NSML: A Machine Learning Platform That Enables You to Focus on Your Models
- Benchmarking State-of-the-Art Deep Learning Software Tools
- The Evolution of Distributed Systems for Graph Neural Networks and their Origin in Graph Processing and Deep Learning: A Survey
- Representation Learning for Dynamic Graphs: A Survey
- Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads
- A newcomer's guide to deep learning for inverse design in nano-photonics
- A robust anomaly finder based on autoencoders
- The Sockeye 2 Neural Machine Translation Toolkit at AMTA 2020
- Communication optimization strategies for distributed deep neural network training: A survey
- Asynchronous Decentralized Parallel Stochastic Gradient Descent
- Scanner: Efficient Video Analysis at Scale
- TensorLayer: A Versatile Library for Efficient Deep Learning Development
- Improving Outfit Recommendation with Co-supervision of Fashion Generation
- Deep Factors for Forecasting
- ChainerMN: Scalable Distributed Deep Learning Framework
- Just ASK: Building an Architecture for Extensible Self-Service Spoken Language Understanding
- TBD: Benchmarking and Analyzing Deep Neural Network Training
- Deep Learning Training in Facebook Data Centers: Design of Scale-up and Scale-out Systems
- Real-time Semantic Image Segmentation via Spatial Sparsity
- Neural Network Distiller: A Python Package For DNN Compression Research
- TensorFlow Eager: A Multi-Stage, Python-Embedded DSL for Machine Learning
- STN-OCR: A single Neural Network for Text Detection and Text Recognition
- pCAMP: Performance Comparison of Machine Learning Packages on the Edges
- Domain Adaptation for Semantic Segmentation via Class-Balanced Self-Training
- Optimizing Network Performance for Distributed DNN Training on GPU Clusters: ImageNet/AlexNet Training in 1.5 Minutes
- Omnivore: An Optimizer for Multi-device Deep Learning on CPUs and GPUs
- An Iterative Machine-Learning Framework for RANS Turbulence Modeling
- TensorOpt: Exploring the Tradeoffs in Distributed DNN Training with Auto-Parallelism
- Adversarial Examples - A Complete Characterisation of the Phenomenon
- Scalable Deep Learning on Distributed Infrastructures: Challenges, Techniques and Tools
- Revise Saturated Activation Functions
- LFFD: A Light and Fast Face Detector for Edge Devices
- Combination of Multiple Global Descriptors for Image Retrieval
- DeepMarks: A Digital Fingerprinting Framework for Deep Neural Networks
- Priority-based Parameter Propagation for Distributed DNN Training
- Attention-guided Chained Context Aggregation for Semantic Segmentation
- Video Big Data Analytics in the Cloud: A Reference Architecture, Survey, Opportunities, and Open Research Issues
- Fusion-Catalyzed Pruning for Optimizing Deep Learning on Intelligent Edge Devices
- Advbox: a toolbox to generate adversarial examples that fool neural networks
- Extensible Structure-Informed Prediction of Formation Energy with Improved Accuracy and Usability employing Neural Networks
- Piecewise Linear Neural Networks and Deep Learning
- HP-GNN: Generating High Throughput GNN Training Implementation on CPU-FPGA Heterogeneous Platform
- AdaScale: Towards Real-time Video Object Detection Using Adaptive Scaling
- Tracklet Association Tracker: An End-to-End Learning-based Association Approach for Multi-Object Tracking
- Semi-Dynamic Load Balancing: Efficient Distributed Learning in Non-Dedicated Environments
- Natural-Parameter Networks: A Class of Probabilistic Neural Networks
- Looking for the Devil in the Details: Learning Trilinear Attention Sampling Network for Fine-grained Image Recognition
- LIFT: Reinforcement Learning in Computer Systems by Learning From Demonstrations
- A Model-Driven Approach to Machine Learning and Software Modeling for the IoT
- PyPose: A Library for Robot Learning with Physics-based Optimization
- Reinforced Genetic Algorithm Learning for Optimizing Computation Graphs
- Why Self-Attention? A Targeted Evaluation of Neural Machine Translation Architectures
- DGL-KE: Training Knowledge Graph Embeddings at Scale
- Mitigate Bias in Face Recognition using Skewness-Aware Reinforcement Learning
- FastPose: Towards Real-time Pose Estimation and Tracking via Scale-normalized Multi-task Networks
- Chester: A Web Delivered Locally Computed Chest X-Ray Disease Prediction System
- Deep Convolutional Neural Networks with Merge-and-Run Mappings
- OneFlow: Redesign the Distributed Deep Learning Framework from Scratch
- Declarative Recursive Computation on an RDBMS, or, Why You Should Use a Database For Distributed Machine Learning
- Confidence Regularized Self-Training
- The Neural Network Approach to Inverse Problems in Differential Equations
- Moniqua: Modulo Quantized Communication in Decentralized SGD
- FedHybrid: A Hybrid Primal-Dual Algorithm Framework for Federated Optimization
- One-Trial Correction of Legacy AI Systems and Stochastic Separation Theorems
- An Efficient Statistical-based Gradient Compression Technique for Distributed Training Systems
- Multi-Objective De Novo Drug Design with Conditional Graph Generative Model
- Sharing Residual Units Through Collective Tensor Factorization in Deep Neural Networks
- Exascale Deep Learning for Scientific Inverse Problems
- SSAP: Single-Shot Instance Segmentation With Affinity Pyramid
- Exploring Object Relation in Mean Teacher for Cross-Domain Detection
- Data-Driven Sparse Structure Selection for Deep Neural Networks
- A Double Residual Compression Algorithm for Efficient Distributed Learning
- Action Machine: Rethinking Action Recognition in Trimmed Videos
- A Good Practice Towards Top Performance of Face Recognition: Transferred Deep Feature Fusion
- Ludwig: a type-based declarative deep learning toolbox
- Binary Neural Networks for Memory-Efficient and Effective Visual Place Recognition in Changing Environments
- AdaCos: Adaptively Scaling Cosine Logits for Effectively Learning Deep Face Representations
- IGCV: Interleaved Structured Sparse Convolutional Neural Networks
- Chainer: A Deep Learning Framework for Accelerating the Research Cycle
- Online Machine Learning in Big Data Streams
- A Survey on Deep Learning Toolkits and Libraries for Intelligent User Interfaces
- Structured Attentions for Visual Question Answering
- Deep Leakage from Gradients
- Fused Text Segmentation Networks for Multi-oriented Scene Text Detection
- A Performance Comparison of Loss Functions for Deep Face Recognition
- Achieving Better Kinship Recognition Through Better Baseline
- TensorDIMM: A Practical Near-Memory Processing Architecture for Embeddings and Tensor Operations in Deep Learning
- TF-Coder: Program Synthesis for Tensor Manipulations
- Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep Learning
- Optimal Subarchitecture Extraction For BERT
- Hemingway: Modeling Distributed Optimization Algorithms
- DL2: A Deep Learning-driven Scheduler for Deep Learning Clusters
- DarkRank: Accelerating Deep Metric Learning via Cross Sample Similarities Transfer
- Local AdaAlter: Communication-Efficient Stochastic Gradient Descent with Adaptive Learning Rates
- MRI Cross-Modality NeuroImage-to-NeuroImage Translation
- LOTR: Face Landmark Localization Using Localization Transformer
- TF-Replicator: Distributed Machine Learning for Researchers
- A Programmable Approach to Neural Network Compression
- Gradient Diversity: a Key Ingredient for Scalable Distributed Learning
- HeAT -- a Distributed and GPU-accelerated Tensor Framework for Data Analytics
- Characterizing the Deep Neural Networks Inference Performance of Mobile Applications
- CIAN: Cross-Image Affinity Net for Weakly Supervised Semantic Segmentation
- Online normalizer calculation for softmax
- Efficient Memory Management for GPU-based Deep Learning Systems
- Parle: parallelizing stochastic gradient descent
- MXNET-MPI: Embedding MPI parallelism in Parameter Server Task Model for scaling Deep Learning
- Probabilistic Forecasting with Temporal Convolutional Neural Network
- SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle
- An Introduction to Deep Learning for the Physical Layer
- Deep Factors with Gaussian Processes for Forecasting
- Dynamic Parameter Allocation in Parameter Servers
- FanStore: Enabling Efficient and Scalable I/O for Distributed Deep Learning
- HAWQV3: Dyadic Neural Network Quantization
- Efficient Execution of Quantized Deep Learning Models: A Compiler Approach
- TensorFlow Estimators: Managing Simplicity vs. Flexibility in High-Level Machine Learning Frameworks
- Seeing Small Faces from Robust Anchor's Perspective
- Regularizing Proxies with Multi-Adversarial Training for Unsupervised Domain-Adaptive Semantic Segmentation
- From Facial Expression Recognition to Interpersonal Relation Prediction
- Graph-Based Global Reasoning Networks
- Integral Human Pose Regression
- Knowledge Projection for Deep Neural Networks
- Performance Modeling and Evaluation of Distributed Deep Learning Frameworks on GPUs
- Deep learning for photoacoustic imaging: a survey
- Structured Binary Neural Networks for Accurate Image Classification and Semantic Segmentation
- nGraph-HE: A Graph Compiler for Deep Learning on Homomorphically Encrypted Data
- Learning Neural Network Subspaces
- Multiple Adaptive Bayesian Linear Regression for Scalable Bayesian Optimization with Warm Start
- Real-Time Machine Learning: The Missing Pieces
- Generative Low-bitwidth Data Free Quantization
- FaceX-Zoo: A PyTorch Toolbox for Face Recognition
- Exploring the Design Space of Deep Convolutional Neural Networks at Large Scale
- MirBot: A collaborative object recognition system for smartphones using convolutional neural networks
- A Labeling-Free Approach to Supervising Deep Neural Networks for Retinal Blood Vessel Segmentation
- Dynamic Mini-batch SGD for Elastic Distributed Training: Learning in the Limbo of Resources
- Memory Warps for Learning Long-Term Online Video Representations
- Learning to play the Chess Variant Crazyhouse above World Champion Level with Deep Neural Networks and Human Data
- Salus: Fine-Grained GPU Sharing Primitives for Deep Learning Applications
- Explicit Interaction Model towards Text Classification
- -Nets: Double Attention Networks
- TedNet: A Pytorch Toolkit for Tensor Decomposition Networks
- The Convergence of Sparsified Gradient Methods
- Multi-Fiber Networks for Video Recognition
- Skin Lesion Analysis Towards Melanoma Detection via End-to-end Deep Learning of Convolutional Neural Networks
- OpenEI: An Open Framework for Edge Intelligence
- Reinforcement Learning and Adaptive Sampling for Optimized DNN Compilation
- Scheduling Computation Graphs of Deep Learning Models on Manycore CPUs
- Optimizing Task Placement and Online Scheduling for Distributed GNN Training Acceleration
- Spectral Feature Transformation for Person Re-identification
- FusionStitching: Boosting Memory Intensive Computations for Deep Learning Workloads
- An Empirical Study towards Characterizing Deep Learning Development and Deployment across Different Frameworks and Platforms
- CROSSBOW: Scaling Deep Learning with Small Batch Sizes on Multi-GPU Servers
- Demystifying Differentiable Programming: Shift/Reset the Penultimate Backpropagator
- Domain-specific Communication Optimization for Distributed DNN Training
- A Learning-based Framework for Hybrid Depth-from-Defocus and Stereo Matching
- Ripple: A Practical Declarative Programming Framework for Serverless Compute
- SparCML: High-Performance Sparse Communication for Machine Learning
- Graph Generative Models for Fast Detector Simulations in High Energy Physics
- DynaComm: Accelerating Distributed CNN Training between Edges and Clouds through Dynamic Communication Scheduling
- cltorch: a Hardware-Agnostic Backend for the Torch Deep Neural Network Library, Based on OpenCL
- Vector representations of text data in deep learning
- Deep Frequent Spatial Temporal Learning for Face Anti-Spoofing
- A Distributed Multi-GPU System for Large-Scale Node Embedding at Tencent
- SMAUG: End-to-End Full-Stack Simulation Infrastructure for Deep Learning Workloads
- Improving Neural Network Training using Dynamic Learning Rate Schedule for PINNs and Image Classification
- Tensor Contraction Layers for Parsimonious Deep Nets
- Training Group Orthogonal Neural Networks with Privileged Information
- KAISA: An Adaptive Second-Order Optimizer Framework for Deep Neural Networks
- Learning Where to Focus for Efficient Video Object Detection
- BMXNet: An Open-Source Binary Neural Network Implementation Based on MXNet
- Deep Hashing with Category Mask for Fast Video Retrieval
- Object Detection in Video with Spatial-temporal Context Aggregation
- Intermittent Demand Forecasting with Deep Renewal Processes
- A survey on Kornia: an Open Source Differentiable Computer Vision Library for PyTorch
- PRETZEL: Opening the Black Box of Machine Learning Prediction Serving Systems
- The Design and Implementation of a Scalable DL Benchmarking Platform
- Parallel Training of Deep Networks with Local Updates
- Improving the Expressiveness of Deep Learning Frameworks with Recursion
- Performance Analysis and Comparison of Distributed Machine Learning Systems
- A Generalized Zero-Shot Quantization of Deep Convolutional Neural Networks via Learned Weights Statistics
- ModelHub.AI: Dissemination Platform for Deep Learning Models
- Distributed Training Large-Scale Deep Architectures
- Face Detection with Feature Pyramids and Landmarks
- Challenges in Migrating Imperative Deep Learning Programs to Graph Execution: An Empirical Study
- Double Supervised Network with Attention Mechanism for Scene Text Recognition
- Unbiased Evaluation of Deep Metric Learning Algorithms
- Performance Evaluation of Deep Learning Tools in Docker Containers
- Neural Machine Translation: A Review of Methods, Resources, and Tools
- Deep Learning Based Computed Tomography Whys and Wherefores
- Question Type Guided Attention in Visual Question Answering
- A DAG Model of Synchronous Stochastic Gradient Descent in Distributed Deep Learning
- Heterogeneity-Aware Asynchronous Decentralized Training
- Data Analytics and Machine Learning Methods, Techniques and Tool for Model-Driven Engineering of Smart IoT Services
- Frustrated with Replicating Claims of a Shared Model? A Solution
- MNN: A Universal and Efficient Inference Engine
- A Framework for Democratizing AI
- daBNN: A Super Fast Inference Framework for Binary Neural Networks on ARM devices
- Progressive Neural Networks for Image Classification
- Deep-Edge: An Efficient Framework for Deep Learning Model Update on Heterogeneous Edge
- N-Adaptive Ritz Method: A Neural Network Enriched Partition of Unity for Boundary Value Problems
- A Proof of Useful Work for Artificial Intelligence on the Blockchain
- Midwifery Learning and Forecasting: Predicting Content Demand with User-Generated Logs
- Auto-STGCN: Autonomous Spatial-Temporal Graph Convolutional Network Search Based on Reinforcement Learning and Existing Research Results
- QR and LQ Decomposition Matrix Backpropagation Algorithms for Square, Wide, and Deep -- Real or Complex -- Matrices and Their Software Implementation
- Parallax: Sparsity-aware Data Parallel Training of Deep Neural Networks
- DeLS-3D: Deep Localization and Segmentation with a 3D Semantic Map
- Echo: Compiler-based GPU Memory Footprint Reduction for LSTM RNN Training
- HierTrain: Fast Hierarchical Edge AI Learning with Hybrid Parallelism in Mobile-Edge-Cloud Computing
- dMath: A Scalable Linear Algebra and Math Library for Heterogeneous GP-GPU Architectures
- Fifer: Tackling Underutilization in the Serverless Era
- Directional Statistics-based Deep Metric Learning for Image Classification and Retrieval
- Ivy: Templated Deep Learning for Inter-Framework Portability
- Diverse Sample Generation: Pushing the Limit of Generative Data-free Quantization
- A Survey on Deep Learning for Neuroimaging-based Brain Disorder Analysis
- A Highly Configurable Hardware/Software Stack for DNN Inference Acceleration
- A Block-wise, Asynchronous and Distributed ADMM Algorithm for General Form Consensus Optimization
- DMLO: Deep Matching LiDAR Odometry
- Optimizing DNN Compilation for Distributed Training with Joint OP and Tensor Fusion
- AOGNets: Compositional Grammatical Architectures for Deep Learning
- Horizontally Fused Training Array: An Effective Hardware Utilization Squeezer for Training Novel Deep Learning Models
- Deep Extreme Multi-label Learning
- Demystifying the MLPerf Benchmark Suite
- Seesaw-Net: Convolution Neural Network With Uneven Group Convolution
- Attentional Feature Fusion
- 3D Context Enhanced Region-based Convolutional Neural Network for End-to-End Lesion Detection
- AI Enabling Technologies: A Survey
- Sionnx: Automatic Unit Test Generator for ONNX Conformance
- The Effectiveness of Discretization in Forecasting: An Empirical Study on Neural Time Series Models
- GraphTheta: A Distributed Graph Neural Network Learning System With Flexible Training Strategy
- Automatic Horizontal Fusion for GPU Kernels
- Auto-MAP: A DQN Framework for Exploring Distributed Execution Plans for DNN Workloads
- Characterizing Deep Learning Training Workloads on Alibaba-PAI
- Adaptive Precision Training: Quantify Back Propagation in Neural Networks with Fixed-point Numbers
- LPRNet: Lightweight Deep Network by Low-rank Pointwise Residual Convolution
- Hybrid Data-Model Parallel Training for Sequence-to-Sequence Recurrent Neural Network Machine Translation
- Hybrid Composition with IdleBlock: More Efficient Networks for Image Recognition
- Efficient Memory Management for Deep Neural Net Inference
- Deep learning in bioinformatics: introduction, application, and perspective in big data era
- Complexity-Weighted Loss and Diverse Reranking for Sentence Simplification
- Fast and Accurate, Convolutional Neural Network Based Approach for Object Detection from UAV
- A Network Structure to Explicitly Reduce Confusion Errors in Semantic Segmentation
- Towards Multi-class Object Detection in Unconstrained Remote Sensing Imagery
- Monocular Depth Estimation with Augmented Ordinal Depth Relationships
- MLtuner: System Support for Automatic Machine Learning Tuning
- A comprehensive study of batch construction strategies for recurrent neural networks in MXNet
- Detection and Attention: Diagnosing Pulmonary Lung Cancer from CT by Imitating Physicians
- Bridging the Gap Between Neural Networks and Neuromorphic Hardware with A Neural Network Compiler
- Theano-MPI: a Theano-based Distributed Training Framework
- RandomOut: Using a convolutional gradient norm to rescue convolutional filters
- Skin disease diagnosis with deep learning: a review
- Distributed Machine Learning through Heterogeneous Edge Systems
- EagerPy: Writing Code That Works Natively with PyTorch, TensorFlow, JAX, and NumPy
- Deep3D: Fully Automatic 2D-to-3D Video Conversion with Deep Convolutional Neural Networks
- ASAP: Asynchronous Approximate Data-Parallel Computation
- Asynchronous Stochastic Proximal Methods for Nonconvex Nonsmooth Optimization
- PydMobileNet: Improved Version of MobileNets with Pyramid Depthwise Separable Convolution
- DISC: A Dynamic Shape Compiler for Machine Learning Workloads
- A Runtime-Based Computational Performance Predictor for Deep Neural Network Training
- DISH: A Distributed Hybrid Primal-Dual Optimization Framework to Utilize System Heterogeneity
- Stanza: Layer Separation for Distributed Training in Deep Learning
- SLSGD: Secure and Efficient Distributed On-device Machine Learning
- Dynamic Stale Synchronous Parallel Distributed Training for Deep Learning
- Faster Distributed Deep Net Training: Computation and Communication Decoupled Stochastic Gradient Descent
- A multi-task convolutional neural network for mega-city analysis using very high resolution satellite imagery and geospatial data
- Stochastic Distributed Optimization for Machine Learning from Decentralized Features
- Machine Learning Automation Toolbox (MLaut)
- Learning Deep Representations Using Convolutional Auto-encoders with Symmetric Skip Connections
- Factorized Bilinear Models for Image Recognition
- Semantic Hierarchy Preserving Deep Hashing for Large-scale Image Retrieval
- swCaffe: a Parallel Framework for Accelerating Deep Learning Applications on Sunway TaihuLight
- Deep Neural Network Assisted Iterative Reconstruction Method for Low Dose CT
- BackPACK: Packing more into backprop
- Optimal Gradient Checkpoint Search for Arbitrary Computation Graphs
- Heterogeneity-aware Gradient Coding for Straggler Tolerance
- A First Look at Deep Learning Apps on Smartphones
- Efficient Training of Convolutional Neural Nets on Large Distributed Systems
- Bootstrapping NLU Models with Multi-task Learning
- An Efficient Method for Face Quality Assessment on the Edge
- SuperNeurons: FFT-based Gradient Sparsification in the Distributed Training of Deep Neural Networks
- Fast, Better Training Trick -- Random Gradient
- Probabilistic Hierarchical Forecasting with Deep Poisson Mixtures
- Deformable Tube Network for Action Detection in Videos
- Network-accelerated Distributed Machine Learning Using MLFabric
- Revisit Batch Normalization: New Understanding from an Optimization View and a Refinement via Composition Optimization
- Daydream: Accurately Estimating the Efficacy of Optimizations for DNN Training
- MetaTune: Meta-Learning Based Cost Model for Fast and Efficient Auto-tuning Frameworks
- Apache Submarine: A Unified Machine Learning Platform Made Simple
- Orchestrate: Infrastructure for Enabling Parallelism during Hyperparameter Optimization
- FastEstimator: A Deep Learning Library for Fast Prototyping and Productization
- ModiPick: SLA-aware Accuracy Optimization For Mobile Deep Inference
- DASNet: Dynamic Activation Sparsity for Neural Network Efficiency Improvement
- Boosting Model Performance through Differentially Private Model Aggregation
- Gaussian Vector: An Efficient Solution for Facial Landmark Detection
- ENT-DESC: Entity Description Generation by Exploring Knowledge Graph
- RLgraph: Modular Computation Graphs for Deep Reinforcement Learning
- A Survey on Large-scale Machine Learning
- Face Detection with End-to-End Integration of a ConvNet and a 3D Model
- Adaptive Periodic Averaging: A Practical Approach to Reducing Communication in Distributed Learning
- Hyperparameter Transfer Learning with Adaptive Complexity
- More Information Supervised Probabilistic Deep Face Embedding Learning
- On Improving Temporal Consistency for Online Face Liveness Detection
- Toward Model Parallelism for Deep Neural Network based on Gradient-free ADMM Framework
- Looking for change? Roll the Dice and demand Attention
- Graph-to-Sequence Learning using Gated Graph Neural Networks
- RPC Considered Harmful: Fast Distributed Deep Learning on RDMA
- A Scalable and Cloud-Native Hyperparameter Tuning System
- Asynch-SGBDT: Asynchronous Parallel Stochastic Gradient Boosting Decision Tree based on Parameters Server
- An Efficient DP-SGD Mechanism for Large Scale NLP Models
- Diagonalwise Refactorization: An Efficient Training Method for Depthwise Convolutions
- Deep Learning to Address Candidate Generation and Cold Start Challenges in Recommender Systems: A Research Survey
- Delving Deep into Liver Focal Lesion Detection: A Preliminary Study
- Parallel Programming Models for Heterogeneous Many-Cores : A Survey
- GEVO: GPU Code Optimization using Evolutionary Computation
- SWIFT: Expedited Failure Recovery for Large-scale DNN Training
- Data-driven forecasting of solar irradiance
- Deep Eyes: Binocular Depth-from-Focus on Focal Stack Pairs
- Optimizer Fusion: Efficient Training with Better Locality and Parallelism
- Adaptive Elastic Training for Sparse Deep Learning on Heterogeneous Multi-GPU Servers
- Efficient and Scalable View Generation from a Single Image using Fully Convolutional Networks
- Luandri: a Clean Lua Interface to the Indri Search Engine
- Draw your Neural Networks
- Comparing the costs of abstraction for DL frameworks
- Dynamic Key-Value Memory Networks for Knowledge Tracing
- Sub-pixel face landmarks using heatmaps and a bag of tricks
- Small, Accurate, and Fast Vehicle Re-ID on the Edge: the SAFR Approach
- Neural Networks for Lorenz Map Prediction: A Trip Through Time
- Elastic deep learning in multi-tenant GPU cluster
- Scheduling Optimization Techniques for Neural Network Training
- Data-parallel distributed training of very large models beyond GPU capacity
- A Survey on Proactive Customer Care: Enabling Science and Steps to Realize it
- SoftNeuro: Fast Deep Inference using Multi-platform Optimization
- Slot Based Image Augmentation System for Object Detection
- A weakly supervised adaptive triplet loss for deep metric learning
- Deep Learning At Scale and At Ease
- Cavs: A Vertex-centric Programming Interface for Dynamic Neural Networks
- Vanishing Nodes: Another Phenomenon That Makes Training Deep Neural Networks Difficult
- DaSGD: Squeezing SGD Parallelization Performance in Distributed Training Using Delayed Averaging
- DEED: A General Quantization Scheme for Communication Efficiency in Bits
- Shuffle-Exchange Brings Faster: Reduce the Idle Time During Communication for Decentralized Neural Network Training
- Training a Binary Weight Object Detector by Knowledge Transfer for Autonomous Driving
- Light-weighted Saliency Detection with Distinctively Lower Memory Cost and Model Size
- GPU-Accelerated Primal Learning for Extremely Fast Large-Scale Classification
- Toward Efficient Online Scheduling for Distributed Machine Learning Systems
- Integrating Deep Learning in Domain Sciences at Exascale
- HypLL: The Hyperbolic Learning Library
- Deep Class-Wise Hashing: Semantics-Preserving Hashing via Class-wise Loss
- Training on the Edge: The why and the how
- Lightweight Mask R-CNN for Long-Range Wireless Power Transfer Systems
- ByzShield: An Efficient and Robust System for Distributed Training
- VERAM: View-Enhanced Recurrent Attention Model for 3D Shape Classification
- Focal Loss Dense Detector for Vehicle Surveillance
- OD-SGD: One-step Delay Stochastic Gradient Descent for Distributed Training
- DBLFace: Domain-Based Labels for NIR-VIS Heterogeneous Face Recognition
- Stochastic Gradient MCMC with Stale Gradients
- Attention as Activation
- Context-Aware Drive-thru Recommendation Service at Fast Food Restaurants
- GPU-based Parallel Computation Support for Stan
- PZnet: Efficient 3D ConvNet Inference on Manycore CPUs
- Large e-retailer image dataset for visual search and product classification
- Good Intentions: Adaptive Parameter Management via Intent Signaling
- DeepLocalize: Fault Localization for Deep Neural Networks
- Beyond the Memory Wall: A Case for Memory-centric HPC System for Deep Learning
- An Accurate Model for Predicting the (Graded) Effect of Context in Word Similarity Based on Bert
- Adaptive Online Learning with Momentum for Contingency-based Voltage Stability Assessment
- Every Filter Extracts A Specific Texture In Convolutional Neural Networks
- Elastic Bulk Synchronous Parallel Model for Distributed Deep Learning
- Towards Quantized Model Parallelism for Graph-Augmented MLPs Based on Gradient-Free ADMM Framework
- ScaleFreeCTR: MixCache-based Distributed Training System for CTR Models with Huge Embedding Table
- Efficient Distributed Semi-Supervised Learning using Stochastic Regularization over Affinity Graphs
- TensorX: Extensible API for Neural Network Model Design and Deployment
- Understanding Neural Machine Translation by Simplification: The Case of Encoder-free Models
- FusionStitching: Boosting Execution Efficiency of Memory Intensive Computations for DL Workloads
- Symbolic Techniques for Deep Learning: Challenges and Opportunities
- Protection of an information system by artificial intelligence: a three-phase approach based on behaviour analysis to detect a hostile scenario
- A Novel Co-design Peta-scale Heterogeneous Cluster for Deep Learning Training
- An Adaptive Remote Stochastic Gradient Method for Training Neural Networks
- How to Train your DNN: The Network Operator Edition
- Dragon: A Computation Graph Virtual Machine Based Deep Learning Framework
- Progressive Compressed Records: Taking a Byte out of Deep Learning Data
- AsymmNet: Towards ultralight convolution neural networks using asymmetrical bottlenecks
- Densely Connected Graph Convolutional Networks for Graph-to-Sequence Learning
- Using Python for Model Inference in Deep Learning
- Sinan: Data-Driven, QoS-Aware Cluster Management for Microservices
- A Survey on Serverless Computing
- Effective GPU Sharing Under Compiler Guidance
- Retention Time of Peptides in Liquid Chromatography Is Well Estimated upon Deep Transfer Learning
- Escaping the abstraction: a foreign function interface for the Unified Form Language [UFL]
- Towards Designing a Self-Managed Machine Learning Inference Serving System inPublic Cloud
- MergeComp: A Compression Scheduler for Scalable Communication-Efficient Distributed Training
- Forget the Learning Rate, Decay Loss
- Variational Bayes Neural Network: Posterior Consistency, Classification Accuracy and Computational Challenges
- Wide Aspect Ratio Matching for Robust Face Detection
- Akid: A Library for Neural Network Research and Production from a Dataism Approach
- Oscars: Adaptive Semi-Synchronous Parallel Model for Distributed Deep Learning with Global View
- Invasiveness Prediction of Pulmonary Adenocarcinomas Using Deep Feature Fusion Networks
- Learning Cascaded Siamese Networks for High Performance Visual Tracking
- Identify the stiffness of DNA via deep learning
- Lightweight, Dynamic Graph Convolutional Networks for AMR-to-Text Generation
- Graph Deep Factors for Forecasting
- LAIF: AI, Deep Learning for Germany Suetterlin Letter Recognition and Generation
- Woodpecker-DL: Accelerating Deep Neural Networks via Hardware-Aware Multifaceted Optimizations
- A temporal-to-spatial deep convolutional neural network for classification of hand movements from multichannel electromyography data
- Using Intuition from Empirical Properties to Simplify Adversarial Training Defense
- A Model Parallel Proximal Stochastic Gradient Algorithm for Partially Asynchronous Systems