Deep Residual Learning for Image Recognition
arXiv:1512.03385
Abstract
Deeper neural networks are more difficult to train. We present a residual learning framework to ease the training of networks that are substantially deeper than those used previously. We explicitly reformulate the layers as learning residual functions with reference to the layer inputs, instead of learning unreferenced functions. We provide comprehensive empirical evidence showing that these residual networks are easier to optimize, and can gain accuracy from considerably increased depth. On the ImageNet dataset we evaluate residual nets with a depth of up to 152 layers---8x deeper than VGG nets but still having lower complexity. An ensemble of these residual nets achieves 3.57% error on the ImageNet test set. This result won the 1st place on the ILSVRC 2015 classification task. We also present analysis on CIFAR-10 with 100 and 1000 layers. The depth of representations is of central importance for many visual recognition tasks. Solely due to our extremely deep representations, we obtain a 28% relative improvement on the COCO object detection dataset. Deep residual nets are foundations of our submissions to ILSVRC & COCO 2015 competitions, where we also won the 1st places on the tasks of ImageNet detection, ImageNet localization, COCO detection, and COCO segmentation.
Tech report
References in corpus (3)
Cited by in corpus (1528)
- A Survey on Deep Learning in Medical Image Analysis
- TensorFlow: A system for large-scale machine learning
- YOLOv3: An Incremental Improvement
- Convolutional Sequence to Sequence Learning
- Attention Mechanisms in Computer Vision: A Survey
- ResUNet-a: a deep learning framework for semantic segmentation of remotely sensed data
- Focal Loss for Dense Object Detection
- Riemannian Walk for Incremental Learning: Understanding Forgetting and Intransigence
- Mixed Precision Training
- Representation Learning on Graphs with Jumping Knowledge Networks
- YouTube-8M: A Large-Scale Video Classification Benchmark
- Contrastive Representation Learning: A Framework and Review
- Snorkel: Rapid Training Data Creation with Weak Supervision
- Searching for Activation Functions
- Efficient Neural Architecture Search via Parameter Sharing
- TernausNet: U-Net with VGG11 Encoder Pre-Trained on ImageNet for Image Segmentation
- AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty
- Achieving Human Parity on Automatic Chinese to English News Translation
- Deep Representation Learning with Part Loss for Person Re-Identification
- COVID-ResNet: A Deep Learning Framework for Screening of COVID19 from Radiographs
- From Perception to Decision: A Data-driven Approach to End-to-end Motion Planning for Autonomous Ground Robots
- Born Again Neural Networks
- Deformable Convolutional Networks
- Decomposing Motion and Content for Natural Video Sequence Prediction
- Deep Predictive Coding Networks for Video Prediction and Unsupervised Learning
- Improved Denoising Diffusion Probabilistic Models
- SphereFace: Deep Hypersphere Embedding for Face Recognition
- Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
- Modeling the Dynamics of PDE Systems with Physics-Constrained Deep Auto-Regressive Networks
- Like What You Like: Knowledge Distill via Neuron Selectivity Transfer
- The Microsoft 2016 Conversational Speech Recognition System
- A Fully Convolutional Neural Network for Cardiac Segmentation in Short-Axis MRI
- #Exploration: A Study of Count-Based Exploration for Deep Reinforcement Learning
- VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection
- Residual Attention Network for Image Classification
- Deep Roots: Improving CNN Efficiency with Hierarchical Filter Groups
- Pyramid Scene Parsing Network
- What makes ImageNet good for transfer learning?
- FusionNet: A deep fully residual convolutional neural network for image segmentation in connectomics
- LocalViT: Analyzing Locality in Vision Transformers
- End-to-End Learning of Geometry and Context for Deep Stereo Regression
- XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks
- CosFace: Large Margin Cosine Loss for Deep Face Recognition
- Real-time Hand Tracking under Occlusion from an Egocentric RGB-D Sensor
- Compact Convolutional Neural Networks for Classification of Asynchronous Steady-state Visual Evoked Potentials
- DeeperGCN: All You Need to Train Deeper GCNs
- Lossy Image Compression with Compressive Autoencoders
- Deep Learning for Physical Processes: Incorporating Prior Scientific Knowledge
- Deep Learning Enables Automatic Detection and Segmentation of Brain Metastases on Multi-Sequence MRI
- Face Mask Detection using Transfer Learning of InceptionV3
- Model compression via distillation and quantization
- Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks
- Deep learning to achieve clinically applicable segmentation of head and neck anatomy for radiotherapy
- Var-CNN: A Data-Efficient Website Fingerprinting Attack Based on Deep Learning
- Residual Dense Network for Image Super-Resolution
- Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions
- Learnable pooling with Context Gating for video classification
- Single Path One-Shot Neural Architecture Search with Uniform Sampling
- Wider or Deeper: Revisiting the ResNet Model for Visual Recognition
- Scalability in Perception for Autonomous Driving: Waymo Open Dataset
- Hierarchical Representations for Efficient Architecture Search
- Superhuman Accuracy on the SNEMI3D Connectomics Challenge
- Star-galaxy Classification Using Deep Convolutional Neural Networks
- Deep Transfer Learning for Person Re-identification
- A Closer Look at Spatiotemporal Convolutions for Action Recognition
- Semantic Instance Segmentation via Deep Metric Learning
- Rapid Prediction of Electron-Ionization Mass Spectrometry using Neural Networks
- You Only Look Twice: Rapid Multi-Scale Object Detection In Satellite Imagery
- Deep Spatio-Temporal Residual Networks for Citywide Crowd Flows Prediction
- Transparency by Design: Closing the Gap Between Performance and Interpretability in Visual Reasoning
- Deep Aesthetic Quality Assessment with Semantic Information
- Attention Enriched Deep Learning Model for Breast Tumor Segmentation in Ultrasound Images
- Sensor-based Gait Parameter Extraction with Deep Convolutional Neural Networks
- Mitigating Adversarial Effects Through Randomization
- Removal of Batch Effects using Distribution-Matching Residual Networks
- CMU DeepLens: Deep Learning For Automatic Image-based Galaxy-Galaxy Strong Lens Finding
- Harmonious Attention Network for Person Re-Identification
- Temporal 3D ConvNets: New Architecture and Transfer Learning for Video Classification
- Natural Adversarial Examples
- An Uncertainty-aware Transfer Learning-based Framework for Covid-19 Diagnosis
- Fathom: Reference Workloads for Modern Deep Learning Methods
- Deformable ConvNets v2: More Deformable, Better Results
- WRPN: Wide Reduced-Precision Networks
- Scaling Out-of-Distribution Detection for Real-World Settings
- Parallel-Data-Free Voice Conversion Using Cycle-Consistent Adversarial Networks
- Pose Invariant Embedding for Deep Person Re-identification
- Glitch Classification and Clustering for LIGO with Deep Transfer Learning
- Deep Transfer Learning: A new deep learning glitch classification method for advanced LIGO
- Dilated Residual Networks
- Classification of Urban Morphology with Deep Learning: Application on Urban Vitality
- CvT: Introducing Convolutions to Vision Transformers
- Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations
- Deep Learning Based Robot for Automatically Picking up Garbage on the Grass
- Visual Entailment: A Novel Task for Fine-Grained Image Understanding
- Identification of heavy, energetic, hadronically decaying particles using machine-learning techniques
- Intrinsic dimension of data representations in deep neural networks
- Cosmological constraints with deep learning from KiDS-450 weak lensing maps
- StegNet: Mega Image Steganography Capacity with Deep Convolutional Network
- Learning Sparse Neural Networks through Regularization
- MONAS: Multi-Objective Neural Architecture Search using Reinforcement Learning
- MDMMT: Multidomain Multimodal Transformer for Video Retrieval
- Examining the Impact of Blur on Recognition by Convolutional Networks
- No bad local minima: Data independent training error guarantees for multilayer neural networks
- Bag of Tricks for Image Classification with Convolutional Neural Networks
- Maximum Classifier Discrepancy for Unsupervised Domain Adaptation
- Improvements to context based self-supervised learning
- Spinal cord gray matter segmentation using deep dilated convolutions
- Machine Learning for the Zwicky Transient Facility
- Deep Learning for Automatic Pneumonia Detection
- Fast Bayesian Optimization of Machine Learning Hyperparameters on Large Datasets
- ImageNet pre-trained models with batch normalization
- An Empirical Model of Large-Batch Training
- Faster CryptoNets: Leveraging Sparsity for Real-World Encrypted Inference
- Learning the PE Header, Malware Detection with Minimal Domain Knowledge
- Multi-graph convolutional network for short-term passenger flow forecasting in urban rail transit
- Image Classification of Melanoma, Nevus and Seborrheic Keratosis by Deep Neural Network Ensemble
- ALWANN: Automatic Layer-Wise Approximation of Deep Neural Network Accelerators without Retraining
- R-C3D: Region Convolutional 3D Network for Temporal Activity Detection
- Beyond Face Rotation: Global and Local Perception GAN for Photorealistic and Identity Preserving Frontal View Synthesis
- SVDNet for Pedestrian Retrieval
- SCNN: An Accelerator for Compressed-sparse Convolutional Neural Networks
- BiSeNet: Bilateral Segmentation Network for Real-time Semantic Segmentation
- Neural Networks for Entity Matching: A Survey
- Compression Artifacts Removal Using Convolutional Neural Networks
- Swapout: Learning an ensemble of deep architectures
- Learning Robust and High-Precision Quantum Controls
- 2018 Robotic Scene Segmentation Challenge
- Deep learning predictions of galaxy merger stage and the importance of observational realism
- Adapting Mask-RCNN for Automatic Nucleus Segmentation
- Amulet: Aggregating Multi-level Convolutional Features for Salient Object Detection
- Multi-scale fully convolutional neural networks for histopathology image segmentation: from nuclear aberrations to the global tissue architecture
- Revisiting Deep Learning Models for Tabular Data
- Face Detection through Scale-Friendly Deep Convolutional Networks
- Training with Quantization Noise for Extreme Model Compression
- Chemception: A Deep Neural Network with Minimal Chemistry Knowledge Matches the Performance of Expert-developed QSAR/QSPR Models
- Train Sparsely, Generate Densely: Memory-efficient Unsupervised Training of High-resolution Temporal GAN
- DeeperCut: A Deeper, Stronger, and Faster Multi-Person Pose Estimation Model
- Multi-style Generative Network for Real-time Transfer
- The Importance of Skip Connections in Biomedical Image Segmentation
- Single-Shot Refinement Neural Network for Object Detection
- Convolutional Neural Networks Applied to Neutrino Events in a Liquid Argon Time Projection Chamber
- B-CNN: Branch Convolutional Neural Network for Hierarchical Classification
- Wavelet Convolutional Neural Networks
- Boosted EfficientNet: Detection of Lymph Node Metastases in Breast Cancer Using Convolutional Neural Network
- Machine learning assisted multiscale modeling of composite phase change materials for Li-ion batteries thermal management
- NSML: Meet the MLaaS platform with a real-world case study
- Benchmarking Neural Network Robustness to Common Corruptions and Surface Variations
- Adaptive Inference through Early-Exit Networks: Design, Challenges and Directions
- Re-ranking Person Re-identification with k-reciprocal Encoding
- Do Deep Convolutional Nets Really Need to be Deep and Convolutional?
- ThiNet: A Filter Level Pruning Method for Deep Neural Network Compression
- Flow-Guided Feature Aggregation for Video Object Detection
- Relay: A New IR for Machine Learning Frameworks
- DeepSense: A Unified Deep Learning Framework for Time-Series Mobile Sensing Data Processing
- Sharpness-Aware Minimization for Efficiently Improving Generalization
- SLEEPNET: Automated Sleep Staging System via Deep Learning
- Finding Strong Gravitational Lenses in the DESI DECam Legacy Survey
- CLIP-Art: Contrastive Pre-training for Fine-Grained Art Classification
- Circle Loss: A Unified Perspective of Pair Similarity Optimization
- Ristretto: Hardware-Oriented Approximation of Convolutional Neural Networks
- What Do We Understand About Convolutional Networks?
- Frame- and Segment-Level Features and Candidate Pool Evaluation for Video Caption Generation
- DeepCorn: A Semi-Supervised Deep Learning Method for High-Throughput Image-Based Corn Kernel Counting and Yield Estimation
- Rethinking ImageNet Pre-training
- Retina U-Net: Embarrassingly Simple Exploitation of Segmentation Supervision for Medical Object Detection
- Fast Online Object Tracking and Segmentation: A Unifying Approach
- Label-Only Membership Inference Attacks
- SCAN: Self-and-Collaborative Attention Network for Video Person Re-identification
- The Shattered Gradients Problem: If resnets are the answer, then what is the question?
- CondenseNet: An Efficient DenseNet using Learned Group Convolutions
- Weakly Supervised Medical Diagnosis and Localization from Multiple Resolutions
- A simple yet effective baseline for 3d human pose estimation
- Machine learning active-nematic hydrodynamics
- DenResCov-19: A deep transfer learning network for robust automatic classification of COVID-19, pneumonia, and tuberculosis from X-rays
- DSOD: Learning Deeply Supervised Object Detectors from Scratch
- Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
- Beyond 512 Tokens: Siamese Multi-depth Transformer-based Hierarchical Encoder for Long-Form Document Matching
- Re-ID done right: towards good practices for person re-identification
- Universal adversarial perturbations
- Learning Deep Context-aware Features over Body and Latent Parts for Person Re-identification
- MAttNet: Modular Attention Network for Referring Expression Comprehension
- Learning a Discriminative Feature Network for Semantic Segmentation
- Towards Interpretable Deep Neural Networks by Leveraging Adversarial Examples
- Learning Longer-term Dependencies in RNNs with Auxiliary Losses
- Measuring the Algorithmic Efficiency of Neural Networks
- Sampling Matters in Deep Embedding Learning
- Predicting Lung Nodule Malignancies by Combining Deep Convolutional Neural Network and Handcrafted Features
- Deep Image Retrieval: Learning global representations for image search
- Recognizing Partial Biometric Patterns
- Autoregressive Convolutional Neural Networks for Asynchronous Time Series
- Peephole: Predicting Network Performance Before Training
- Deep Koalarization: Image Colorization using CNNs and Inception-ResNet-v2
- Incorporating Symmetry into Deep Dynamics Models for Improved Generalization
- Learning Rich Features for Image Manipulation Detection
- Orthographic Feature Transform for Monocular 3D Object Detection
- Instance-aware Semantic Segmentation via Multi-task Network Cascades
- Generative Adversarial Residual Pairwise Networks for One Shot Learning
- ME R-CNN: Multi-Expert R-CNN for Object Detection
- Interleaved Group Convolutions for Deep Neural Networks
- The Devil is in the Middle: Exploiting Mid-level Representations for Cross-Domain Instance Matching
- Dissecting the Graphcore IPU Architecture via Microbenchmarking
- Learning Spatio-Temporal Features with 3D Residual Networks for Action Recognition
- APE-GAN: Adversarial Perturbation Elimination with GAN
- Putting An End to End-to-End: Gradient-Isolated Learning of Representations
- Efficient and Robust Parallel DNN Training through Model Parallelism on Multi-GPU Platform
- Preserving Semantic Relations for Zero-Shot Learning
- New high-quality strong lens candidates with deep learning in the Kilo Degree Survey
- Complex-YOLO: Real-time 3D Object Detection on Point Clouds
- A Low Effort Approach to Structured CNN Design Using PCA
- GAN-based Virtual Re-Staining: A Promising Solution for Whole Slide Image Analysis
- Adversarial Complementary Learning for Weakly Supervised Object Localization
- Neural Message Passing with Edge Updates for Predicting Properties of Molecules and Materials
- Gated Siamese Convolutional Neural Network Architecture for Human Re-Identification
- Residual Connections Encourage Iterative Inference
- Image-Based Size Analysis of Agglomerated and Partially Sintered Particles via Convolutional Neural Networks
- TDAN: Temporally Deformable Alignment Network for Video Super-Resolution
- Dynamic Few-Shot Visual Learning without Forgetting
- UNIT-DDPM: UNpaired Image Translation with Denoising Diffusion Probabilistic Models
- The unreasonable effectiveness of the forget gate
- Full Page Handwriting Recognition via Image to Sequence Extraction
- OpenEDS: Open Eye Dataset
- Learning Deep Models for Face Anti-Spoofing: Binary or Auxiliary Supervision
- CNN-SLAM: Real-time dense monocular SLAM with learned depth prediction
- Illuminating Pedestrians via Simultaneous Detection & Segmentation
- Incorporating the Knowledge of Dermatologists to Convolutional Neural Networks for the Diagnosis of Skin Lesions
- EfficientPose: An efficient, accurate and scalable end-to-end 6D multi object pose estimation approach
- Discriminability objective for training descriptive captions
- The One Hundred Layers Tiramisu: Fully Convolutional DenseNets for Semantic Segmentation
- Towards 3D Human Pose Estimation in the Wild: a Weakly-supervised Approach
- CityPersons: A Diverse Dataset for Pedestrian Detection
- SPCANet: Stellar Parameters and Chemical Abundances Network for LAMOST-II Medium Resolution Survey
- AddNet: Deep Neural Networks Using FPGA-Optimized Multipliers
- Latent Weights Do Not Exist: Rethinking Binarized Neural Network Optimization
- Semantic-aware Grad-GAN for Virtual-to-Real Urban Scene Adaption
- Co-training for Demographic Classification Using Deep Learning from Label Proportions
- Pyramid Stereo Matching Network
- Low-Shot Learning from Imaginary Data
- Stacking-Based Deep Neural Network: Deep Analytic Network for Pattern Classification
- Privacy-preserving Machine Learning through Data Obfuscation
- And the Bit Goes Down: Revisiting the Quantization of Neural Networks
- Person Search with Natural Language Description
- Improved Recurrent Neural Networks for Session-based Recommendations
- How Can We Be So Dense? The Benefits of Using Highly Sparse Representations
- VideoLSTM Convolves, Attends and Flows for Action Recognition
- Three-Dimensional Shapes of Spinning Helium Nanodroplets
- A Real-time Hand Gesture Recognition and Human-Computer Interaction System
- Deep Unsupervised Saliency Detection: A Multiple Noisy Labeling Perspective
- Intriguing Properties of Contrastive Losses
- Memory-Efficient Pipeline-Parallel DNN Training
- Semantic Tagging with Deep Residual Networks
- COVID-19 Monitoring System using Social Distancing and Face Mask Detection on Surveillance video datasets
- Deep Reinforcement Learning-based Image Captioning with Embedding Reward
- Learning Feature Pyramids for Human Pose Estimation
- Visual Translation Embedding Network for Visual Relation Detection
- Adversarial PoseNet: A Structure-aware Convolutional Network for Human Pose Estimation
- Cyclic Functional Mapping: Self-supervised correspondence between non-isometric deformable shapes
- CDC: Convolutional-De-Convolutional Networks for Precise Temporal Action Localization in Untrimmed Videos
- Strong lens systems search in the Dark Energy Survey using Convolutional Neural Networks
- Posterior Network: Uncertainty Estimation without OOD Samples via Density-Based Pseudo-Counts
- High-quality strong lens candidates in the final Kilo Degree survey footprint
- Detecting and Recognizing Human-Object Interactions
- Attack-Resistant Federated Learning with Residual-based Reweighting
- TFPose: Direct Human Pose Estimation with Transformers
- Improving Fast Segmentation With Teacher-student Learning
- Deep Transfer Learning for Star Cluster Classification: I. Application to the PHANGS-HST Survey
- Comparison of Batch Normalization and Weight Normalization Algorithms for the Large-scale Image Classification
- Morphological classification of galaxies with deep learning: comparing 3-way and 4-way CNNs
- Reducing Noise in GAN Training with Variance Reduced Extragradient
- FermiNets: Learning generative machines to generate efficient neural networks via generative synthesis
- Mosaic Flows: A Transferable Deep Learning Framework for Solving PDEs on Unseen Domains
- PointFusion: Deep Sensor Fusion for 3D Bounding Box Estimation
- Lets keep it simple, Using simple architectures to outperform deeper and more complex architectures
- Cascade Residual Learning: A Two-stage Convolutional Neural Network for Stereo Matching
- Attentive Generative Adversarial Network for Raindrop Removal from a Single Image
- DeepStreaks: identifying fast-moving objects in the Zwicky Transient Facility data with deep learning
- Deblending galaxy superpositions with branched generative adversarial networks
- DeepSource: Point Source Detection using Deep Learning
- Image-Image Domain Adaptation with Preserved Self-Similarity and Domain-Dissimilarity for Person Re-identification
- Square Kilometre Array Science Data Challenge 1: analysis and results
- Fast and Accurate Image Super-Resolution with Deep Laplacian Pyramid Networks
- Deep Variation-structured Reinforcement Learning for Visual Relationship and Attribute Detection
- A survey of biodiversity informatics: Concepts, practices, and challenges
- Deep residual detection of radio frequency interference for FAST
- TIC 168789840: A Sextuply-Eclipsing Sextuple Star System
- Stiffness: A New Perspective on Generalization in Neural Networks
- Diversity Regularized Spatiotemporal Attention for Video-based Person Re-identification
- A volumetric deep Convolutional Neural Network for simulation of mock dark matter halo catalogues
- A Comparison of CNN-based Face and Head Detectors for Real-Time Video Surveillance Applications
- PiCANet: Learning Pixel-wise Contextual Attention for Saliency Detection
- Deep SCNN-based Real-time Object Detection for Self-driving Vehicles Using LiDAR Temporal Data
- SS-CAM: Smoothed Score-CAM for Sharper Visual Feature Localization
- Probabilistic Simulation of Quantum Circuits with the Transformer
- MoleculeNet: A Benchmark for Molecular Machine Learning
- Pix3D: Dataset and Methods for Single-Image 3D Shape Modeling
- Dynamic Edge-Conditioned Filters in Convolutional Neural Networks on Graphs
- Boosting Adversarial Attacks with Momentum
- Large-Scale Image Retrieval with Attentive Deep Local Features
- MEC: Memory-efficient Convolution for Deep Neural Network
- Generative Image Translation for Data Augmentation in Colorectal Histopathology Images
- Integrated Object Detection and Tracking with Tracklet-Conditioned Detection
- Emergent Translation in Multi-Agent Communication
- AstroVaDEr: Astronomical Variational Deep Embedder for Unsupervised Morphological Classification of Galaxies and Synthetic Image Generation
- On the effectiveness of task granularity for transfer learning
- SegFlow: Joint Learning for Video Object Segmentation and Optical Flow
- Deep Spatial Feature Reconstruction for Partial Person Re-identification: Alignment-Free Approach
- Self-Supervision Closes the Gap Between Weak and Strong Supervision in Histology
- Efficient Sparse-Winograd Convolutional Neural Networks
- Lorentz Boost Networks: Autonomous Physics-Inspired Feature Engineering
- Deep Learning at Scale for the Construction of Galaxy Catalogs in the Dark Energy Survey
- DeepMerge II: Building Robust Deep Learning Algorithms for Merging Galaxy Identification Across Domains
- WasteNet: Waste Classification at the Edge for Smart Bins
- Comprehensive analysis of a dense sample of FRB 121102 bursts
- Characterising Bias in Compressed Models
- CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
- YouTube-BoundingBoxes: A Large High-Precision Human-Annotated Data Set for Object Detection in Video
- Optimized and autonomous machine learning framework for characterizing pores, particles, grains and grain boundaries in microstructural images
- Dual Attention Networks for Multimodal Reasoning and Matching
- GResNet: Graph Residual Network for Reviving Deep GNNs from Suspended Animation
- Scalable Person Re-identification on Supervised Smoothed Manifold
- CANet: Class-Agnostic Segmentation Networks with Iterative Refinement and Attentive Few-Shot Learning
- Neutrino interaction classification with a convolutional neural network in the DUNE far detector
- Stacked Generative Adversarial Networks
- Using convolutional neural networks to predict galaxy metallicity from three-color images
- Multiple Document Datasets Pre-training Improves Text Line Detection With Deep Neural Networks
- Towards an astronomical foundation model for stars with a Transformer-based model
- Cache Telepathy: Leveraging Shared Resource Attacks to Learn DNN Architectures
- Word2VisualVec: Image and Video to Sentence Matching by Visual Feature Prediction
- Convolutional Neural Nets in Chemical Engineering: Foundations, Computations, and Applications
- EvalAI: Towards Better Evaluation Systems for AI Agents
- On the dissection of degenerate cosmologies with machine learning
- Learning deep structured active contours end-to-end
- On the Feasibility of Generic Deep Disaggregation for Single-Load Extraction
- Data Distillation: Towards Omni-Supervised Learning
- Deep Video Deblurring
- Deep Discriminative Clustering Analysis
- SpotTune: Transfer Learning through Adaptive Fine-tuning
- A deep architecture for unified aesthetic prediction
- An Attention Free Transformer
- Priority-based Parameter Propagation for Distributed DNN Training
- High Performance Zero-Memory Overhead Direct Convolutions
- Mobile Video Object Detection with Temporally-Aware Feature Maps
- Revisiting Video Saliency: A Large-scale Benchmark and a New Model
- SafetyNet: Detecting and Rejecting Adversarial Examples Robustly
- End-to-End Dense Video Captioning with Masked Transformer
- Incorporating External Knowledge to Answer Open-Domain Visual Questions with Dynamic Memory Networks
- Higher Order Recurrent Neural Networks
- From Lost to Found: Discover Missing UI Design Semantics through Recovering Missing Tags
- Annotating Object Instances with a Polygon-RNN
- Combining Background Subtraction Algorithms with Convolutional Neural Network
- Direct Detection of Dark Matter Substructure in Strong Lens Images with Convolutional Neural Networks
- BlockDrop: Dynamic Inference Paths in Residual Networks
- Structured Attention Guided Convolutional Neural Fields for Monocular Depth Estimation
- Ranked Reward: Enabling Self-Play Reinforcement Learning for Combinatorial Optimization
- SCA-CNN: Spatial and Channel-wise Attention in Convolutional Networks for Image Captioning
- Rethinking the Usage of Batch Normalization and Dropout in the Training of Deep Neural Networks
- Understanding the Disharmony between Dropout and Batch Normalization by Variance Shift
- Evaluating prose style transfer with the Bible
- Patch-based Progressive 3D Point Set Upsampling
- Application of Decision Rules for Handling Class Imbalance in Semantic Segmentation
- Stealing Links from Graph Neural Networks
- DeepRecSys: A System for Optimizing End-To-End At-scale Neural Recommendation Inference
- Visual Features for Context-Aware Speech Recognition
- PVN3D: A Deep Point-wise 3D Keypoints Voting Network for 6DoF Pose Estimation
- Weakly- and Semi-Supervised Object Detection with Expectation-Maximization Algorithm
- NASNet: A Neuron Attention Stage-by-Stage Net for Single Image Deraining
- Detecting Out-of-Distribution Inputs in Deep Neural Networks Using an Early-Layer Output
- MaskLab: Instance Segmentation by Refining Object Detection with Semantic and Direction Features
- Towards High Performance Video Object Detection for Mobiles
- Camera Style Adaptation for Person Re-identification
- Document AI: Benchmarks, Models and Applications
- DecomposeMe: Simplifying ConvNets for End-to-End Learning
- Few Shot Speaker Recognition using Deep Neural Networks
- Selective Convolutional Descriptor Aggregation for Fine-Grained Image Retrieval
- Hex2vec -- Context-Aware Embedding H3 Hexagons with OpenStreetMap Tags
- OctNet: Learning Deep 3D Representations at High Resolutions
- Max-Mahalanobis Linear Discriminant Analysis Networks
- Temporal Generative Adversarial Nets with Singular Value Clipping
- Deep Convolutional Neural Networks with Merge-and-Run Mappings
- Robust Quantization: One Model to Rule Them All
- Deep Verifier Networks: Verification of Deep Discriminative Models with Deep Generative Models
- Pulsar Candidate Identification Using Semi-Supervised Generative Adversarial Networks
- A Comprehensive Survey of Machine Learning Applied to Radar Signal Processing
- An Interactive Data Visualization and Analytics Tool to Evaluate Mobility and Sociability Trends During COVID-19
- Learning Better Features for Face Detection with Feature Fusion and Segmentation Supervision
- Zero-shot Knowledge Transfer via Adversarial Belief Matching
- SFD: Single Shot Scale-invariant Face Detector
- All You Need is Beyond a Good Init: Exploring Better Solution for Training Extremely Deep Convolutional Neural Networks with Orthonormality and Modulation
- Oriented Response Networks
- Concurrent Activity Recognition with Multimodal CNN-LSTM Structure
- Dice Loss for Data-imbalanced NLP Tasks
- Understanding Infographics through Textual and Visual Tag Prediction
- Galaxy Morphology Classification using Neural Ordinary Differential Equations
- Learning Deep Neural Networks for Vehicle Re-ID with Visual-spatio-temporal Path Proposals
- PDNet: Semantic Segmentation integrated with a Primal-Dual Network for Document binarization
- Zero-shot Recognition via Semantic Embeddings and Knowledge Graphs
- Deep Triplet Ranking Networks for One-Shot Recognition
- Im2Avatar: Colorful 3D Reconstruction from a Single Image
- Grouped Convolutional Neural Networks for Multivariate Time Series
- It Takes Two to Tango: Towards Theory of AI's Mind
- Convolutional Neural Pyramid for Image Processing
- Quantifying Perceptual Distortion of Adversarial Examples
- Maintaining Discrimination and Fairness in Class Incremental Learning
- Baryon acoustic oscillations reconstruction using convolutional neural networks
- Training Competitive Binary Neural Networks from Scratch
- Show, Adapt and Tell: Adversarial Training of Cross-domain Image Captioner
- Sharing Residual Units Through Collective Tensor Factorization in Deep Neural Networks
- Learning Deep ResNet Blocks Sequentially using Boosting Theory
- FibeR-CNN: Expanding Mask R-CNN to Improve Image-Based Fiber Analysis
- Supervision-by-Registration: An Unsupervised Approach to Improve the Precision of Facial Landmark Detectors
- Bayesian Stokes inversion with Normalizing flows
- Using Deep Learning for Segmentation and Counting within Microscopy Data
- Natural Environment Benchmarks for Reinforcement Learning
- Bright, Relatively Isolated Star Clusters in PHANGS-HST Galaxies: Aperture Corrections, Quantitative Morphologies, and Comparison with Synthetic Stellar Population Models
- Neural ODE and Holographic QCD
- Speaking the Same Language: Matching Machine to Human Captions by Adversarial Training
- Reliable Probability Forecast of Solar Flares: Deep Flare Net-Reliable (DeFN-R)
- TinySpeech: Attention Condensers for Deep Speech Recognition Neural Networks on Edge Devices
- Efficient Continual Learning with Modular Networks and Task-Driven Priors
- Rethinking Class-Balanced Methods for Long-Tailed Visual Recognition from a Domain Adaptation Perspective
- Impression Network for Video Object Detection
- Grappa -- A Machine Learned Molecular Mechanics Force Field
- Deep Pyramidal Residual Networks
- Can Graph Neural Networks Count Substructures?
- Contrastive Learning Inverts the Data Generating Process
- Action Machine: Rethinking Action Recognition in Trimmed Videos
- MegDet: A Large Mini-Batch Object Detector
- Text as Neural Operator: Image Manipulation by Text Instruction
- SipMask: Spatial Information Preservation for Fast Image and Video Instance Segmentation
- Siamese Cascaded Region Proposal Networks for Real-Time Visual Tracking
- Single Image Deraining using Scale-Aware Multi-Stage Recurrent Network
- TFLMS: Large Model Support in TensorFlow by Graph Rewriting
- Well Googled is Half Done: Multimodal Forecasting of New Fashion Product Sales with Image-based Google Trends
- 3D Human Pose Estimation with Relational Networks
- Hyperparameter Optimization: A Spectral Approach
- Deep learning enabled multi-wavelength spatial coherence microscope for the classification of malaria-infected stages with limited labelled data size
- Detecting Photoshopped Faces by Scripting Photoshop
- COVID-19 Screening Using Residual Attention Network an Artificial Intelligence Approach
- Weakly Supervised Dense Video Captioning
- What Do You See? Evaluation of Explainable Artificial Intelligence (XAI) Interpretability through Neural Backdoors
- Shared Data and Algorithms for Deep Learning in Fundamental Physics
- RankIQA: Learning from Rankings for No-reference Image Quality Assessment
- Fashion IQ: A New Dataset Towards Retrieving Images by Natural Language Feedback
- Regularizing RNNs for Caption Generation by Reconstructing The Past with The Present
- DVQA: Understanding Data Visualizations via Question Answering
- Iterative Visual Reasoning Beyond Convolutions
- High-Resolution Image Inpainting using Multi-Scale Neural Patch Synthesis
- Wide Compression: Tensor Ring Nets
- ImageAssist: Tools for Enhancing Touchscreen-Based Image Exploration Systems for Blind and Low Vision Users
- Searches for Compact Binary Coalescence Events using Neural Networks in LIGO/Virgo Second Observation Period
- Learning Affinity via Spatial Propagation Networks
- HARRISON: A Benchmark on HAshtag Recommendation for Real-world Images in Social Networks
- Scribbler: Controlling Deep Image Synthesis with Sketch and Color
- Learning to Detect Human-Object Interactions
- Text Detection and Recognition in the Wild: A Review
- Less Is More: Picking Informative Frames for Video Captioning
- Deep Back-Projection Networks For Super-Resolution
- Pose-Invariant Face Alignment with a Single CNN
- Quark-Gluon Jet Discrimination Using Convolutional Neural Networks
- Interpolated Adversarial Training: Achieving Robust Neural Networks without Sacrificing Too Much Accuracy
- How fine can fine-tuning be? Learning efficient language models
- How Does Batch Normalization Help Binary Training?
- Effect of Annotation Errors on Drone Detection with YOLOv3
- Instance-level Human Parsing via Part Grouping Network
- Machine Learning in High Energy Physics: A review of heavy-flavor jet tagging at the LHC
- Learning Dual Convolutional Neural Networks for Low-Level Vision
- 3D-A-Nets: 3D Deep Dense Descriptor for Volumetric Shapes with Adversarial Networks
- Player Identification in Hockey Broadcast Videos
- Deep Learning for Identifying Iran's Cultural Heritage Buildings in Need of Conservation Using Image Classification and Grad-CAM
- Predictive Analysis of Diabetic Retinopathy with Transfer Learning
- A semi-supervised self-training method to develop assistive intelligence for segmenting multiclass bridge elements from inspection videos
- Online Video Deblurring via Dynamic Temporal Blending Network
- Models Matter, So Does Training: An Empirical Study of CNNs for Optical Flow Estimation
- Improving Multi-Modal Learning with Uni-Modal Teachers
- Demonstration of background rejection using deep convolutional neural networks in the NEXT experiment
- Image to Image Translation for Domain Adaptation
- DDRprog: A CLEVR Differentiable Dynamic Reasoning Programmer
- Quantization for Rapid Deployment of Deep Neural Networks
- DenseNet Models for Tiny ImageNet Classification
- Abnormality Detection in Mammography using Deep Convolutional Neural Networks
- Rethinking Feature Distribution for Loss Functions in Image Classification
- Towards Frequency-Based Explanation for Robust CNN
- CASENet: Deep Category-Aware Semantic Edge Detection
- Estimating 6D Pose From Localizing Designated Surface Keypoints
- Significance-aware Information Bottleneck for Domain Adaptive Semantic Segmentation
- Person Search via A Mask-Guided Two-Stream CNN Model
- Not All Pixels Are Equal: Difficulty-aware Semantic Segmentation via Deep Layer Cascade
- Jet-Parton Assignment in ttH Events using Deep Learning
- ElephantBook: A Semi-Automated Human-in-the-Loop System for Elephant Re-Identification
- NeuPDE: Neural Network Based Ordinary and Partial Differential Equations for Modeling Time-Dependent Data
- Teaching Categories to Human Learners with Visual Explanations
- A Convolutional Neural Network For Cosmic String Detection in CMB Temperature Maps
- Crowd counting via scale-adaptive convolutional neural network
- Learning efficient haptic shape exploration with a rigid tactile sensor array
- Natural Language Guided Visual Relationship Detection
- From phonemes to images: levels of representation in a recurrent neural model of visually-grounded language learning
- Motion-Appearance Co-Memory Networks for Video Question Answering
- BlockCNN: A Deep Network for Artifact Removal and Image Compression
- Interpretable Convolutional Neural Networks
- Acting Thoughts: Towards a Mobile Robotic Service Assistant for Users with Limited Communication Skills
- Colorectal cancer diagnosis from histology images: A comparative study
- Task-driven Visual Saliency and Attention-based Visual Question Answering
- From Patch to Image Segmentation using Fully Convolutional Networks -- Application to Retinal Images
- Solar Image Restoration with the Cycle-GAN Based on Multi-Fractal Properties of Texture Features
- AdaptivFloat: A Floating-point based Data Type for Resilient Deep Learning Inference
- Abdominal synthetic CT reconstruction with intensity projection prior for MRI-only adaptive radiotherapy
- DCAN: Dual Channel-wise Alignment Networks for Unsupervised Scene Adaptation
- A Fast and Accurate One-Stage Approach to Visual Grounding
- A convolutional neural network approach for reconstructing polarization information of photoelectric X-ray polarimeters
- Constrained Neural Ordinary Differential Equations with Stability Guarantees
- Anatomy of Catastrophic Forgetting: Hidden Representations and Task Semantics
- Improving Fully Convolution Network for Semantic Segmentation
- Deep Neural Network Architectures for Modulation Classification
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent Kernel
- Polygonal Building Segmentation by Frame Field Learning
- Revisiting Model Stitching to Compare Neural Representations
- Knowledge Concentration: Learning 100K Object Classifiers in a Single CNN
- SARM: Sparse Autoregressive Model for Scalable Generation of Sparse Images in Particle Physics
- Video Captioning via Hierarchical Reinforcement Learning
- Real-Time Value-Driven Data Augmentation in the Era of LSST
- An Analysis of Visual Question Answering Algorithms
- Viewmaker Networks: Learning Views for Unsupervised Representation Learning
- Situation Recognition with Graph Neural Networks
- Macroscale fracture surface segmentation via semi-supervised learning considering the structural similarity
- High-Resolution Multispectral Dataset for Semantic Segmentation
- Cosine Normalization: Using Cosine Similarity Instead of Dot Product in Neural Networks
- Batch Kalman Normalization: Towards Training Deep Neural Networks with Micro-Batches
- Lattice Long Short-Term Memory for Human Action Recognition
- Image Super-Resolution via Dual-State Recurrent Networks
- All Weather Perception: Joint Data Association, Tracking, and Classification for Autonomous Ground Vehicles
- Network of Experts for Large-Scale Image Categorization
- Pneumothorax Segmentation: Deep Learning Image Segmentation to predict Pneumothorax
- Repulsion Loss: Detecting Pedestrians in a Crowd
- SAUNet: Shape Attentive U-Net for Interpretable Medical Image Segmentation
- Weakly-supervised Visual Grounding of Phrases with Linguistic Structures
- Comparison of projection domain, image domain, and comprehensive deep learning for sparse-view X-ray CT image reconstruction
- Practical Block-wise Neural Network Architecture Generation
- A Closer Look at Local Aggregation Operators in Point Cloud Analysis
- Mass Estimation of Galaxy Clusters with Deep Learning II: CMB Cluster Lensing
- Learning a Discriminative Prior for Blind Image Deblurring
- Quantizing Convolutional Neural Networks for Low-Power High-Throughput Inference Engines
- RUN:Residual U-Net for Computer-Aided Detection of Pulmonary Nodules without Candidate Selection
- Q-CapsNets: A Specialized Framework for Quantizing Capsule Networks
- PPR-FCN: Weakly Supervised Visual Relation Detection via Parallel Pairwise R-FCN
- Zero-Shot Visual Recognition using Semantics-Preserving Adversarial Embedding Networks
- Multiple People Tracking Using Hierarchical Deep Tracklet Re-identification
- An Accurate and Real-time Self-blast Glass Insulator Location Method Based On Faster R-CNN and U-net with Aerial Images
- Pose-Robust Face Recognition via Deep Residual Equivariant Mapping
- Knowledge Adaptation for Efficient Semantic Segmentation
- Stable Tensor Neural Networks for Rapid Deep Learning
- Automatic localization and decoding of honeybee markers using deep convolutional neural networks
- MRI Tumor Segmentation with Densely Connected 3D CNN
- Agile Amulet: Real-Time Salient Object Detection with Contextual Attention
- Automated Quality Control of Vacuum Insulated Glazing by Convolutional Neural Network Image Classification
- End-to-end Concept Word Detection for Video Captioning, Retrieval, and Question Answering
- Take it in your stride: Do we need striding in CNNs?
- A survey of Object Classification and Detection based on 2D/3D data
- Using deep Residual Networks to search for galaxy-Lyα emitter lens candidates based on spectroscopic-selection
- Sketch-R2CNN: An Attentive Network for Vector Sketch Recognition
- Comparison of Deep learning models on time series forecasting : a case study of Dissolved Oxygen Prediction
- Unsupervised Learning of Monocular Depth Estimation with Bundle Adjustment, Super-Resolution and Clip Loss
- Star Cluster Classification using Deep Transfer Learning with PHANGS-HST
- Physical Accuracy of Deep Neural Networks for 2D and 3D Multi-Mineral Segmentation of Rock micro-CT Images
- Suppressing the Unusual: towards Robust CNNs using Symmetric Activation Functions
- SemStyle: Learning to Generate Stylised Image Captions using Unaligned Text
- All about Structure: Adapting Structural Information across Domains for Boosting Semantic Segmentation
- Learning to Segment Instances in Videos with Spatial Propagation Network
- Learning Convolutional Networks for Content-weighted Image Compression
- Scene Graph Generation via Conditional Random Fields
- Semantic Compositional Networks for Visual Captioning
- Consensus-based Sequence Training for Video Captioning
- LabelBank: Revisiting Global Perspectives for Semantic Segmentation
- Temporal Non-Volume Preserving Approach to Facial Age-Progression and Age-Invariant Face Recognition
- Few-Shot Image Recognition by Predicting Parameters from Activations
- Characterizing and Improving Stability in Neural Style Transfer
- Learning Discriminative Motion Features Through Detection
- Unsupervised state representation learning with robotic priors: a robustness benchmark
- Efficient Execution of Quantized Deep Learning Models: A Compiler Approach
- Stabilizing Gradients for Deep Neural Networks via Efficient SVD Parameterization
- Tensorial Neural Networks: Generalization of Neural Networks and Application to Model Compression
- PeerNets: Exploiting Peer Wisdom Against Adversarial Attacks
- Personalized Federated Deep Learning for Pain Estimation From Face Images
- Robust Policies via Mid-Level Visual Representations: An Experimental Study in Manipulation and Navigation
- Towards Optimal Signal Extraction for Imaging X-ray Polarimetry
- A Pixel-Based Framework for Data-Driven Clothing
- Connecting optical morphology, environment, and HI mass fraction for low-redshift galaxies using deep learning
- Fast robust peg-in-hole insertion with continuous visual servoing
- Doubly Attentive Transformer Machine Translation
- Feature Enhancement Network: A Refined Scene Text Detector
- Generating Descriptions with Grounded and Co-Referenced People
- Where to Focus: Query Adaptive Matching for Instance Retrieval Using Convolutional Feature Maps
- Full-Resolution Residual Networks for Semantic Segmentation in Street Scenes
- Improving Simple Models with Confidence Profiles
- Optimized Custom Dataset for Efficient Detection of Underwater Trash
- Learning Digital Camera Pipeline for Extreme Low-Light Imaging
- Materials property prediction using symmetry-labeled graphs as atomic-position independent descriptors
- Seeing Small Faces from Robust Anchor's Perspective
- Spec-ResNet: A General Audio Steganalysis scheme based on Deep Residual Network of Spectrogram
- Appliance Detection Using Very Low-Frequency Smart Meter Time Series
- Multi-Scale Deep Learning for Estimating Horizontal Velocity Fields on the Solar Surface
- Astroconformer: The Prospects of Analyzing Stellar Light Curves with Transformer-Based Deep Learning Models
- Attend to the Difference: Cross-Modality Person Re-identification via Contrastive Correlation
- The iNaturalist Species Classification and Detection Dataset
- Programmable Neural Network Trojan for Pre-Trained Feature Extractor
- Global Planar Convolutions for improved context aggregation in Brain Tumor Segmentation
- CTAP: Complementary Temporal Action Proposal Generation
- Learning to Evaluate Image Captioning
- Decoding Dark Matter Substructure without Supervision
- Monocular 3D Human Pose Estimation In The Wild Using Improved CNN Supervision
- Structured Binary Neural Networks for Accurate Image Classification and Semantic Segmentation
- Graph-Based Global Reasoning Networks
- An Attempt towards Interpretable Audio-Visual Video Captioning
- FacePoseNet: Making a Case for Landmark-Free Face Alignment
- Advancing Self-supervised Monocular Depth Learning with Sparse LiDAR
- Training Deeper Neural Machine Translation Models with Transparent Attention
- Breaking Batch Normalization for better explainability of Deep Neural Networks through Layer-wise Relevance Propagation
- Skeleton Key: Image Captioning by Skeleton-Attribute Decomposition
- Recognition Of Surface Defects On Steel Sheet Using Transfer Learning
- Scene Parsing with Global Context Embedding
- Towards Generalizable Surgical Activity Recognition Using Spatial Temporal Graph Convolutional Networks
- Single View Stereo Matching
- EraseReLU: A Simple Way to Ease the Training of Deep Convolution Neural Networks
- Art of singular vectors and universal adversarial perturbations
- Learning to Synthesize a 4D RGBD Light Field from a Single Image
- Instance Embedding Transfer to Unsupervised Video Object Segmentation
- A deep learning based solution for construction equipment detection: from development to deployment
- Scalable inference with Autoregressive Neural Ratio Estimation
- SuPer Deep: A Surgical Perception Framework for Robotic Tissue Manipulation using Deep Learning for Feature Extraction
- A Smartphone based Application for Skin Cancer Classification Using Deep Learning with Clinical Images and Lesion Information
- Binary Ensemble Neural Network: More Bits per Network or More Networks per Bit?
- Rafiki: Machine Learning as an Analytics Service System
- Adaptive Semantic Segmentation with a Strategic Curriculum of Proxy Labels
- Rethinking Re-Sampling in Imbalanced Semi-Supervised Learning
- PlaneNet: Piece-wise Planar Reconstruction from a Single RGB Image
- S3K: Self-Supervised Semantic Keypoints for Robotic Manipulation via Multi-View Consistency
- Large Scale Language Modeling: Converging on 40GB of Text in Four Hours
- Measuring the Transferability of Adversarial Examples
- Fusion++: Volumetric Object-Level SLAM
- A Read-Write Memory Network for Movie Story Understanding
- KTAN: Knowledge Transfer Adversarial Network
- Iris and periocular recognition in arabian race horses using deep convolutional neural networks
- Learning to Segment Every Thing
- 3DB: A Framework for Debugging Computer Vision Models
- Privacy-Preserving Action Recognition for Smart Hospitals using Low-Resolution Depth Images
- Focal Loss based Residual Convolutional Neural Network for Speech Emotion Recognition
- Set-valued classification -- overview via a unified framework
- Triple consistency loss for pairing distributions in GAN-based face synthesis
- Deep Spatial Regression Model for Image Crowd Counting
- RMM: A Recursive Mental Model for Dialog Navigation
- Clickbait Identification using Neural Networks
- Correlated and Individual Multi-Modal Deep Learning for RGB-D Object Recognition
- Constrained CycleGAN for Effective Generation of Ultrasound Sector Images of Improved Spatial Resolution
- Bayesian Inference with Generative Adversarial Network Priors
- Single-Shot Object Detection with Enriched Semantics
- Exploring the Design Space of Deep Convolutional Neural Networks at Large Scale
- Convolutional Random Walk Networks for Semantic Image Segmentation
- Unsupervised Learning from Continuous Video in a Scalable Predictive Recurrent Network
- AdaStereo: A Simple and Efficient Approach for Adaptive Stereo Matching
- Deep Extreme Cut: From Extreme Points to Object Segmentation
- Learning Visual Knowledge Memory Networks for Visual Question Answering
- Automatic Diagnosis of Pneumothorax from Chest Radiographs: A Systematic Literature Review
- CNN-based Facial Affect Analysis on Mobile Devices
- To Share or Not To Share: A Comprehensive Appraisal of Weight-Sharing
- Learning to Extract a Video Sequence from a Single Motion-Blurred Image
- PolyTransform: Deep Polygon Transformer for Instance Segmentation
- A Revisit on Deep Hashings for Large-scale Content Based Image Retrieval
- AI-assisted super-resolution cosmological simulations III: Time evolution
- On Periodic Functions as Regularizers for Quantization of Neural Networks
- Tell-and-Answer: Towards Explainable Visual Question Answering using Attributes and Captions
- Dynamic backdoor attacks against federated learning
- ScaleNet: Guiding Object Proposal Generation in Supermarkets and Beyond
- Prototype-based Neural Network Layers: Incorporating Vector Quantization
- Fair Comparison: Quantifying Variance in Resultsfor Fine-grained Visual Categorization
- An Empirical Analysis of the Impact of Data Augmentation on Knowledge Distillation
- Improved Part Segmentation Performance by Optimising Realism of Synthetic Images using Cycle Generative Adversarial Networks
- An Analysis of Scale Invariance in Object Detection - SNIP
- Tails: Chasing Comets with the Zwicky Transient Facility and Deep Learning
- Understanding Traffic Density from Large-Scale Web Camera Data
- Joint Prediction of Depths, Normals and Surface Curvature from RGB Images using CNNs
- ML-LBM: Machine Learning Aided Flow Simulation in Porous Media
- Learning Shape Priors for Single-View 3D Completion and Reconstruction
- Noise2Void - Learning Denoising from Single Noisy Images
- Vision Xformers: Efficient Attention for Image Classification
- Relation Networks for Optic Disc and Fovea Localization in Retinal Images
- Reconstruction of IACT events using deep learning techniques with CTLearn
- Spectral Feature Transformation for Person Re-identification
- Multi-scale Location-aware Kernel Representation for Object Detection
- A Multimodal Late Fusion Model for E-Commerce Product Classification
- NILMFormer: Non-Intrusive Load Monitoring that Accounts for Non-Stationarity
- Fast, Exact and Multi-Scale Inference for Semantic Image Segmentation with Deep Gaussian CRFs
- Skip-Clip: Self-Supervised Spatiotemporal Representation Learning by Future Clip Order Ranking
- Attend and Interact: Higher-Order Object Interactions for Video Understanding
- Learning to Navigate Using Mid-Level Visual Priors
- CROSSBOW: Scaling Deep Learning with Small Batch Sizes on Multi-GPU Servers
- UTRNet: High-Resolution Urdu Text Recognition In Printed Documents
- Semi-Supervised Multitask Learning on Multispectral Satellite Images Using Wasserstein Generative Adversarial Networks (GANs) for Predicting Poverty
- Artificial Intelligence Assisted Inversion (AIAI) of Synthetic Type Ia Supernova Spectra
- TAFE-Net: Task-Aware Feature Embeddings for Low Shot Learning
- Deep Heterogeneous Feature Fusion for Template-Based Face Recognition
- Image-based localization using LSTMs for structured feature correlation
- LCNN: Lookup-based Convolutional Neural Network
- Regressing Robust and Discriminative 3D Morphable Models with a very Deep Neural Network
- Modularized Morphing of Neural Networks
- CNN-based event classification for alpha-decay events in nuclear emulsion
- Deep learning the astrometric signature of dark matter substructure
- Coordinating Filters for Faster Deep Neural Networks
- AdvGAN++ : Harnessing latent layers for adversary generation
- Eformer: Edge Enhancement based Transformer for Medical Image Denoising
- Training with the Invisibles: Obfuscating Images to Share Safely for Learning Visual Recognition Models
- Spatial-Temporal Self-Attention Network for Flow Prediction
- Towards Robust RGB-D Human Mesh Recovery
- Pointwise Convolutional Neural Networks
- Combining Image Features and Patient Metadata to Enhance Transfer Learning
- LightTrack: A Generic Framework for Online Top-Down Human Pose Tracking
- cGANs with Multi-Hinge Loss
- SRN: Side-output Residual Network for Object Symmetry Detection in the Wild
- InverseFaceNet: Deep Monocular Inverse Face Rendering
- An End-to-End Approach to Natural Language Object Retrieval via Context-Aware Deep Reinforcement Learning
- DA-GAN: Instance-level Image Translation by Deep Attention Generative Adversarial Networks (with Supplementary Materials)
- Weakly Supervised Attention Model for RV StrainClassification from volumetric CTPA Scans
- Underwater object detection using Invert Multi-Class Adaboost with deep learning
- Single-Shot Bidirectional Pyramid Networks for High-Quality Object Detection
- DSAC - Differentiable RANSAC for Camera Localization
- VIPriors 1: Visual Inductive Priors for Data-Efficient Deep Learning Challenges
- Label-Driven Reconstruction for Domain Adaptation in Semantic Segmentation
- Iterative Deep Learning for Network Topology Extraction
- OffRoadTranSeg: Semi-Supervised Segmentation using Transformers on OffRoad environments
- Real-Time High-Resolution Background Matting
- Deep learning for detection of bird vocalisations
- Deep Restricted Boltzmann Networks
- Improved Stereo Matching with Constant Highway Networks and Reflective Confidence Learning
- Graph neural network for 3D classification of ambiguities and optical crosstalk in scintillator-based neutrino detectors
- Weakly-Supervised 3D Pose Estimation from a Single Image using Multi-View Consistency
- Detection of Einstein Telescope gravitational wave signals from binary black holes using deep learning
- Generative Partition Networks for Multi-Person Pose Estimation
- Gradient Scheduling with Global Momentum for Non-IID Data Distributed Asynchronous Training
- Holistic Interstitial Lung Disease Detection using Deep Convolutional Neural Networks: Multi-label Learning and Unordered Pooling
- NLP-CUET@DravidianLangTech-EACL2021: Investigating Visual and Textual Features to Identify Trolls from Multimodal Social Media Memes
- TIRAMISU: A Polyhedral Compiler for Dense and Sparse Deep Learning
- Deep Texture Manifold for Ground Terrain Recognition
- Reducing Uncertainty in Undersampled MRI Reconstruction with Active Acquisition
- MHP-VOS: Multiple Hypotheses Propagation for Video Object Segmentation
- Signal Combination for Language Identification
- Weak-lensing Mass Reconstruction of Galaxy Clusters with Convolutional Neural Network
- Recurrent Localization Networks applied to the Lippmann-Schwinger Equation
- Robust Processing-In-Memory Neural Networks via Noise-Aware Normalization
- Autonomous Deep Learning: A Genetic DCNN Designer for Image Classification
- Dual Path Networks for Multi-Person Human Pose Estimation
- Towards recognizing the light facet of the Higgs Boson
- SAWNet: A Spatially Aware Deep Neural Network for 3D Point Cloud Processing
- Joint Pruning & Quantization for Extremely Sparse Neural Networks
- An Efficient Transformer Decoder with Compressed Sub-layers
- Pose Guided Human Video Generation
- Domain Alignment with Triplets
- Scaling GRPC Tensorflow on 512 nodes of Cori Supercomputer
- Reconstruction of 3-D Atomic Distortions from Electron Microscopy with Deep Learning
- PhishGAN: Data Augmentation and Identification of Homoglpyh Attacks
- Occupancy Map Prediction Using Generative and Fully Convolutional Networks for Vehicle Navigation
- Neural networks with differentiable structure
- Compiling Deep Learning Models for Custom Hardware Accelerators
- A Generalised Signature Method for Multivariate Time Series Feature Extraction
- Turning a Blind Eye: Explicit Removal of Biases and Variation from Deep Neural Network Embeddings
- How To Extract Fashion Trends From Social Media? A Robust Object Detector With Support For Unsupervised Learning
- Recurrent Filter Learning for Visual Tracking
- Image segmentation via Cellular Automata
- Adaptive Feeding: Achieving Fast and Accurate Detections by Adaptively Combining Object Detectors
- Evaluating Transfer Learning for Simplifying GitHub READMEs
- 2D Image Relighting with Image-to-Image Translation
- DocVQA: A Dataset for VQA on Document Images
- Fetal Gender Identification using Machine and Deep Learning Algorithms on Phonocardiogram Signals
- OrigamiNet: Weakly-Supervised, Segmentation-Free, One-Step, Full Page Text Recognition by learning to unfold
- 3D Reconstruction in Canonical Co-ordinate Space from Arbitrarily Oriented 2D Images
- Deep Learning applied to Road Traffic Speed forecasting
- ActionFlowNet: Learning Motion Representation for Action Recognition
- Initialization Matters: Regularizing Manifold-informed Initialization for Neural Recommendation Systems
- Semi-supervised Domain Adaptation based on Dual-level Domain Mixing for Semantic Segmentation
- Learning Visual Question Answering by Bootstrapping Hard Attention
- Semantic Segmentation for Partially Occluded Apple Trees Based on Deep Learning
- Learning a Discriminative Filter Bank within a CNN for Fine-grained Recognition
- Optimization Algorithm Inspired Deep Neural Network Structure Design
- Self-Supervised Feature Learning by Learning to Spot Artifacts
- Geometric robustness of deep networks: analysis and improvement
- DeepDeblur: Fast one-step blurry face images restoration
- Deep Feature Flow for Video Recognition
- Optimization of Artificial Neural Networks models applied to the identification of images of asteroids' resonant arguments
- Shape-from-Mask: A Deep Learning Based Human Body Shape Reconstruction from Binary Mask Images
- Exploring Corruption Robustness: Inductive Biases in Vision Transformers and MLP-Mixers
- Context-Aware Visual Compatibility Prediction
- A Continuous Convolutional Trainable Filter for Modelling Unstructured Data
- Comprehensive Evaluation of OpenCL-based Convolutional Neural Network Accelerators in Xilinx and Altera FPGAs
- An Empirical Exploration of Skip Connections for Sequential Tagging
- An LSTM-Based Dynamic Customer Model for Fashion Recommendation
- FairyTailor: A Multimodal Generative Framework for Storytelling
- Incremental Sequence Learning
- EC-GAN: Low-Sample Classification using Semi-Supervised Algorithms and GANs
- RarePlanes: Synthetic Data Takes Flight
- OCoR: An Overlapping-Aware Code Retriever
- Channel-wise Hessian Aware trace-Weighted Quantization of Neural Networks
- Knowledge-based Radiation Treatment Planning: A Data-driven Method Survey
- Multimodal Memory Modelling for Video Captioning
- FOTS: Fast Oriented Text Spotting with a Unified Network
- Super-Resolution with Deep Adaptive Image Resampling
- Weaving Multi-scale Context for Single Shot Detector
- Deep learning for semantic segmentation of remote sensing images with rich spectral content
- SketchParse : Towards Rich Descriptions for Poorly Drawn Sketches using Multi-Task Hierarchical Deep Networks
- A Simple Loss Function for Improving the Convergence and Accuracy of Visual Question Answering Models
- Learning to Learn from Noisy Web Videos
- Gated Context Model with Embedded Priors for Deep Image Compression
- Learning Representations of Sets through Optimized Permutations
- Lightweight Convolutional Representations for On-Device Natural Language Processing
- Character-based NMT with Transformer
- FireNet: Real-time Segmentation of Fire Perimeter from Aerial Video
- SCREENet: A Multi-view Deep Convolutional Neural Network for Classification of High-resolution Synthetic Mammographic Screening Scans
- A(DP)SGD: Asynchronous Decentralized Parallel Stochastic Gradient Descent with Differential Privacy
- Advances in Asynchronous Parallel and Distributed Optimization
- A deep learning theory for neural networks grounded in physics
- Efficient Palm-Line Segmentation with U-Net Context Fusion Module
- Scalable, End-to-End, Deep-Learning-Based Data Reconstruction Chain for Particle Imaging Detectors
- Evaluating Soccer Player: from Live Camera to Deep Reinforcement Learning
- Re-identification = Retrieval + Verification: Back to Essence and Forward with a New Metric
- Dual ResGCN for Balanced Scene GraphGeneration
- Conditional Extreme Value Theory for Open Set Video Domain Adaptation
- Foreground Removal of CO Intensity Mapping Using Deep Learning
- COVID-19 Detection in Chest X-ray Images Using Swin-Transformer and Transformer in Transformer
- Soft Calibration Objectives for Neural Networks
- Federated Noisy Client Learning
- Wise-SrNet: A Novel Architecture for Enhancing Image Classification by Learning Spatial Resolution of Feature Maps
- Transform and Tell: Entity-Aware News Image Captioning
- Benanza: Automatic Benchmark Generation to Compute "Lower-bound" Latency and Inform Optimizations of Deep Learning Models on GPUs
- Timage -- A Robust Time Series Classification Pipeline
- Transformable Bottleneck Networks
- Unsupervised Single Image Deraining with Self-supervised Constraints
- Efficient Coarse-to-Fine Non-Local Module for the Detection of Small Objects
- Orders-of-magnitude speedup in atmospheric chemistry modeling through neural network-based emulation
- Detecting Cyberattacks in Industrial Control Systems Using Convolutional Neural Networks
- A Simple Cache Model for Image Recognition
- Question Type Guided Attention in Visual Question Answering
- TOM-Net: Learning Transparent Object Matting from a Single Image
- Training Neural Networks by Using Power Linear Units (PoLUs)
- Self-Supervised Relative Depth Learning for Urban Scene Understanding
- Aesthetic-Driven Image Enhancement by Adversarial Learning
- Building Emotional Machines: Recognizing Image Emotions through Deep Neural Networks
- Deep Blind Image Inpainting
- Towards Understanding the Data Dependency of Mixup-style Training
- Deep Quantization: Encoding Convolutional Activations with Deep Generative Model
- Dense Recurrent Neural Networks for Scene Labeling
- AVA: Adversarial Vignetting Attack against Visual Recognition
- Cosmic Background Removal with Deep Neural Networks in SBND
- Fast Neural Architecture Construction using EnvelopeNets
- Robust Face Detection via Learning Small Faces on Hard Images
- ResFPN: Residual Skip Connections in Multi-Resolution Feature Pyramid Networks for Accurate Dense Pixel Matching
- Rock Classification in Petrographic Thin Section Images Based on Concatenated Convolutional Neural Networks
- DeepGalaxy: Deducing the Properties of Galaxy Mergers from Images Using Deep Neural Networks
- Dual-Glance Model for Deciphering Social Relationships
- Privacy-preserving medical image analysis
- Can we cover navigational perception needs of the visually impaired by panoptic segmentation?
- Railway Anomaly detection model using synthetic defect images generated by CycleGAN
- HateProof: Are Hateful Meme Detection Systems really Robust?
- Estimating Photometric Redshifts for Galaxies from the DESI Legacy Imaging Surveys with Bayesian Neural Networks Trained by DESI EDR
- Inferring Cosmic String Tension through the Neural Network Prediction of String Locations in CMB Maps
- Weakly and Semi Supervised Human Body Part Parsing via Pose-Guided Knowledge Transfer
- A 3D Coarse-to-Fine Framework for Volumetric Medical Image Segmentation
- Transfer entropy-based feedback improves performance in artificial neural networks
- Unmanned Aerial Vehicle Visual Detection and Tracking using Deep Neural Networks: A Performance Benchmark
- LaSO: Label-Set Operations networks for multi-label few-shot learning
- Deep Class Aware Denoising
- Statistically Motivated Second Order Pooling
- UC Merced Submission to the ActivityNet Challenge 2016
- Training Keyword Spotters with Limited and Synthesized Speech Data
- Language-Driven Image Style Transfer
- ToyArchitecture: Unsupervised Learning of Interpretable Models of the World
- Automatic Liver Segmentation from CT Images Using Deep Learning Algorithms: A Comparative Study
- Deep neural network ensemble by data augmentation and bagging for skin lesion classification
- On transfer learning using a MAC model variant
- A Generalized Network for MRI Intensity Normalization
- Neither Private Nor Fair: Impact of Data Imbalance on Utility and Fairness in Differential Privacy
- To Compress, or Not to Compress: Characterizing Deep Learning Model Compression for Embedded Inference
- DAF:re: A Challenging, Crowd-Sourced, Large-Scale, Long-Tailed Dataset For Anime Character Recognition
- Truth or Backpropaganda? An Empirical Investigation of Deep Learning Theory
- Between-class Learning for Image Classification
- Convolutional Regression for Visual Tracking
- Imprinto: Enhancing Infrared Inkjet Watermarking for Human and Machine Perception
- Echo: Compiler-based GPU Memory Footprint Reduction for LSTM RNN Training
- An Empirical Evaluation of Adversarial Robustness under Transfer Learning
- Crowd Transformer Network
- Convolution Aware Initialization
- The Phonexia VoxCeleb Speaker Recognition Challenge 2021 System Description
- Predicting bulge to total luminosity ratio of galaxies using deep learning
- Do Normalization Layers in a Deep ConvNet Really Need to Be Distinct?
- ModuleNet: Knowledge-inherited Neural Architecture Search
- Deep driven fMRI decoding of visual categories
- COLD: Concurrent Loads Disaggregator for Non-Intrusive Load Monitoring
- Zero-Resource Neural Machine Translation with Multi-Agent Communication Game
- Trade-offs in Top-k Classification Accuracies on Losses for Deep Learning
- Stochastic Downsampling for Cost-Adjustable Inference and Improved Regularization in Convolutional Networks
- Deep Dose Plugin Towards Real-time Monte Carlo Dose Calculation Through a Deep Learning based Denoising Algorithm
- DenseImage Network: Video Spatial-Temporal Evolution Encoding and Understanding
- Self-labeled Conditional GANs
- Depth-Based 3D Hand Pose Estimation: From Current Achievements to Future Goals
- Deep Stacked Networks with Residual Polishing for Image Inpainting
- A Theory of Multiple-Source Adaptation with Limited Target Labeled Data
- Faster Convergence in Deep-Predictive-Coding Networks to Learn Deeper Representations
- GradSign: Model Performance Inference with Theoretical Insights
- Towards Structured Analysis of Broadcast Badminton Videos
- Phonology Recognition in American Sign Language
- Nonuniform-to-Uniform Quantization: Towards Accurate Quantization via Generalized Straight-Through Estimation
- Deep Learning for Metagenomic Data: using 2D Embeddings and Convolutional Neural Networks
- Towards High Performance Video Object Detection
- DeepSkeleton: Skeleton Map for 3D Human Pose Regression
- A Machine Learning Approach to Correcting Atmospheric Seeing in Solar Flare Observations
- Anomalous Sound Detection as a Simple Binary Classification Problem with Careful Selection of Proxy Outlier Examples
- YOLOStereo3D: A Step Back to 2D for Efficient Stereo 3D Detection
- Learning 3D Shapes as Multi-Layered Height-maps using 2D Convolutional Networks
- Demystifying the MLPerf Benchmark Suite
- End-to-End Deep Kronecker-Product Matching for Person Re-identification
- SegStereo: Exploiting Semantic Information for Disparity Estimation
- Domain-Specific Priors and Meta Learning for Few-Shot First-Person Action Recognition
- Learning Frequency-aware Dynamic Network for Efficient Super-Resolution
- Neural Network Training as an Optimal Control Problem: An Augmented Lagrangian Approach
- Towards large-scale, automated, accurate detection of CCTV camera objects using computer vision. Applications and implications for privacy, safety, and cybersecurity. (Preprint)
- Joint Discovery of Object States and Manipulation Actions
- Tearing Down the Memory Wall
- Visual Forecasting by Imitating Dynamics in Natural Sequences
- Imitating Targets from all sides: An Unsupervised Transfer Learning method for Person Re-identification
- Do Vision Models Encode Object-Level Semantic Relatedness? A Cognitive Psychology-Inspired Benchmark
- RILOD: Near Real-Time Incremental Learning for Object Detection at the Edge
- FaceSpoof Buster: a Presentation Attack Detector Based on Intrinsic Image Properties and Deep Learning
- Rethinking Training from Scratch for Object Detection
- A flexible FPGA accelerator for convolutional neural networks
- UV-GAN: Adversarial Facial UV Map Completion for Pose-invariant Face Recognition
- Local Deep Implicit Functions for 3D Shape
- Beyond Forward Shortcuts: Fully Convolutional Master-Slave Networks (MSNets) with Backward Skip Connections for Semantic Segmentation
- SPATL: Salient Parameter Aggregation and Transfer Learning for Heterogeneous Clients in Federated Learning
- Video Frame Interpolation by Plug-and-Play Deep Locally Linear Embedding
- Irregular Convolutional Neural Networks
- toon2real: Translating Cartoon Images to Realistic Images
- Pano2CAD: Room Layout From A Single Panorama Image
- One-Shot Speaker Identification for a Service Robot using a CNN-based Generic Verifier
- EvoGrad: Efficient Gradient-Based Meta-Learning and Hyperparameter Optimization
- Diagnosing Error in Temporal Action Detectors
- Semi-supervised Skin Detection by Network with Mutual Guidance
- How to Make a BLT Sandwich? Learning to Reason towards Understanding Web Instructional Videos
- A morphological segmentation approach to determining bar lengths
- ADF & TransApp: A Transformer-Based Framework for Appliance Detection Using Smart Meter Consumption Series
- Fast Spectral Ranking for Similarity Search
- BusyHands: A Hand-Tool Interaction Database for Assembly Tasks Semantic Segmentation
- Deep neural network for solving differential equations motivated by Legendre-Galerkin approximation
- Learning Inertial Odometry for Dynamic Legged Robot State Estimation
- Training Deep Neural Networks Without Batch Normalization
- Theano-MPI: a Theano-based Distributed Training Framework
- Non-Negative Bregman Divergence Minimization for Deep Direct Density Ratio Estimation
- Noisy Computations during Inference: Harmful or Helpful?
- Simulating Personal Food Consumption Patterns using a Modified Markov Chain
- Biased Bytes: On the Validity of Estimating Food Consumption from Digital Traces
- MotionSqueeze: Neural Motion Feature Learning for Video Understanding
- Scene Understanding Networks for Autonomous Driving based on Around View Monitoring System
- Deep Residual Network for Joint Demosaicing and Super-Resolution
- Video Object Segmentation with Joint Re-identification and Attention-Aware Mask Propagation
- Extension of Direct Feedback Alignment to Convolutional and Recurrent Neural Network for Bio-plausible Deep Learning
- Inter-Patient ECG Classification with Convolutional and Recurrent Neural Networks
- Assessing Intelligence in Artificial Neural Networks
- CNN Acceleration by Low-rank Approximation with Quantized Factors
- HyPar-Flow: Exploiting MPI and Keras for Scalable Hybrid-Parallel DNN Training using TensorFlow
- 3D Object Detection From LiDAR Data Using Distance Dependent Feature Extraction
- Learning Normalized Inputs for Iterative Estimation in Medical Image Segmentation
- Food Ingredients Recognition through Multi-label Learning
- Towards Zero-shot Cross-lingual Image Retrieval and Tagging
- Doubly Nested Network for Resource-Efficient Inference
- MedSelect: Selective Labeling for Medical Image Classification Combining Meta-Learning with Deep Reinforcement Learning
- Measuring Fairness in Generative Models
- Dual Encoder Fusion U-Net (DEFU-Net) for Cross-manufacturer Chest X-ray Segmentation
- Convolutional Neural Network Interpretability with General Pattern Theory
- A Winning Hand: Compressing Deep Networks Can Improve Out-Of-Distribution Robustness
- Facilitating Access to Multilingual COVID-19 Information via Neural Machine Translation
- Active Perception for Ambiguous Objects Classification
- Matrix and tensor decompositions for training binary neural networks
- An Effective Pipeline for a Real-world Clothes Retrieval System
- Why Do Deep Neural Networks Still Not Recognize These Images?: A Qualitative Analysis on Failure Cases of ImageNet Classification
- Darker than Black-Box: Face Reconstruction from Similarity Queries
- Elastic Gossip: Distributing Neural Network Training Using Gossip-like Protocols
- Measuring the Substructure Mass Power Spectrum of 23 SLACS Strong Galaxy-Galaxy Lenses with Convolutional Neural Networks
- The Evolution of Neural Network-Based Chart Patterns: A Preliminary Study
- fruit-SALAD: A Style Aligned Artwork Dataset to reveal similarity perception in image embeddings
- Residual Frames with Efficient Pseudo-3D CNN for Human Action Recognition
- HySTER: A Hybrid Spatio-Temporal Event Reasoner
- OAH-Net: A Deep Neural Network for Hologram Reconstruction of Off-axis Digital Holographic Microscope
- CNN aided Weighted Interpolation for Channel Estimation in Vehicular Communications
- A Cascaded Residual UNET for Fully Automated Segmentation of Prostate and Peripheral Zone in T2-weighted 3D Fast Spin Echo Images
- Beyond Trade-off: Accelerate FCN-based Face Detector with Higher Accuracy
- Application-Driven Near-Data Processing for Similarity Search
- Predicting Localized Primordial Star Formation with Deep Convolutional Neural Networks
- Low-Cost Parameterizations of Deep Convolutional Neural Networks
- Towards Label-Free 3D Segmentation of Optical Coherence Tomography Images of the Optic Nerve Head Using Deep Learning
- Ultrasound-Guided Robotic Navigation with Deep Reinforcement Learning
- Neural Language Priors
- 3D-DETNet: a Single Stage Video-Based Vehicle Detector
- Towards Human-Machine Cooperation: Self-supervised Sample Mining for Object Detection
- Deep Flow-Guided Video Inpainting
- Learning to Support: Exploiting Structure Information in Support Sets for One-Shot Learning
- Unsupervised Domain Adaptation for Spatio-Temporal Action Localization
- Unified Generator-Classifier for Efficient Zero-Shot Learning
- Compression of Deep Neural Networks for Image Instance Retrieval
- Identity Preserving Generative Adversarial Network for Cross-Domain Person Re-identification
- Deep Network Interpolation for Continuous Imagery Effect Transition
- Extreme 3D Face Reconstruction: Seeing Through Occlusions
- SeGAN: Segmenting and Generating the Invisible
- Flexible and Scalable Deep Learning with MMLSpark
- DF-Net: Unsupervised Joint Learning of Depth and Flow using Cross-Task Consistency
- Semi-Supervised Noisy Student Pre-training on EfficientNet Architectures for Plant Pathology Classification
- Constraint Solving with Deep Learning for Symbolic Execution
- Identifying Most Walkable Direction for Navigation in an Outdoor Environment
- Deep-learning-based identification of odontogenic keratocysts in hematoxylin- and eosin-stained jaw cyst specimens
- EdgeSegNet: A Compact Network for Semantic Segmentation
- Multi-domain semantic segmentation with pyramidal fusion
- DelugeNets: Deep Networks with Efficient and Flexible Cross-layer Information Inflows
- A Multi-view RGB-D Approach for Human Pose Estimation in Operating Rooms
- Evolving Deep Neural Networks by Multi-objective Particle Swarm Optimization for Image Classification
- A polarization analyzer from Deep Neural Networks
- Multi-Expert Gender Classification on Age Group by Integrating Deep Neural Networks
- Visualizing the decision-making process in deep neural decision forest
- OkwuGbé: End-to-End Speech Recognition for Fon and Igbo
- Diagnostic Image Quality Assessment and Classification in Medical Imaging: Opportunities and Challenges
- Structure Learning of Deep Networks via DNA Computing Algorithm
- FrameNet: Learning Local Canonical Frames of 3D Surfaces from a Single RGB Image
- MHAttnSurv: Multi-Head Attention for Survival Prediction Using Whole-Slide Pathology Images
- LDMNet: Low Dimensional Manifold Regularized Neural Networks
- Hardware Acceleration of Explainable Machine Learning using Tensor Processing Units
- TACo: Token-aware Cascade Contrastive Learning for Video-Text Alignment
- Classifying CMB time-ordered data through deep neural networks
- Unsupervised Video Depth Estimation Based on Ego-motion and Disparity Consensus
- Unconstrained optimisation on Riemannian manifolds
- Gradient Information Guided Deraining with A Novel Network and Adversarial Training
- Spontaneous Symmetry Breaking in Neural Networks
- EfficientHRNet: Efficient Scaling for Lightweight High-Resolution Multi-Person Pose Estimation
- Energy-efficient Amortized Inference with Cascaded Deep Classifiers
- How to train accurate BNNs for embedded systems?
- Opening the Blackbox: Accelerating Neural Differential Equations by Regularizing Internal Solver Heuristics
- Cluster Analysis with Deep Embeddings and Contrastive Learning
- Real-time regression analysis with deep convolutional neural networks
- Loss Barcode: A Topological Measure of Escapability in Loss Landscapes
- Face Aging with Contextual Generative Adversarial Nets
- Exponential Moving Average Model in Parallel Speech Recognition Training
- Cost-Effective Training of Deep CNNs with Active Model Adaptation
- Low-Latency Video Semantic Segmentation
- Grounding inductive biases in natural images:invariance stems from variations in data
- Classification of simulated radio signals using Wide Residual Networks for use in the search for extra-terrestrial intelligence
- AI4AI: Quantitative Methods for Classifying Host Species from Avian Influenza DNA Sequence
- A Self-supervised Approach for Adversarial Robustness
- TensorFlow with user friendly Graphical Framework for object detection API
- Multi-Miner: Object-Adaptive Region Mining for Weakly-Supervised Semantic Segmentation
- StarMap for Category-Agnostic Keypoint and Viewpoint Estimation
- A model for interpreting social interactions in local image regions
- Butterfly Effect: Bidirectional Control of Classification Performance by Small Additive Perturbation
- Search to Distill: Pearls are Everywhere but not the Eyes
- Frustum VoxNet for 3D object detection from RGB-D or Depth images
- Synthetic Unknown Class Learning for Learning Unknowns
- Three-Stream Convolutional Networks for Video-based Person Re-Identification
- Neural Style Representations and the Large-Scale Classification of Artistic Style
- Efficient Training of Convolutional Neural Nets on Large Distributed Systems
- Learnable Adaptive Cosine Estimator (LACE) for Image Classification
- Generating Nontrivial Melodies for Music as a Service
- Fast Marching Energy CNN
- Face Recognition in Unconstrained Conditions: A Systematic Review
- Modular Learning Component Attacks: Today's Reality, Tomorrow's Challenge
- Design Environment of Quantization-Aware Edge AI Hardware for Few-Shot Learning
- Weight Map Layer for Noise and Adversarial Attack Robustness
- Tracking Without Re-recognition in Humans and Machines
- SurfConv: Bridging 3D and 2D Convolution for RGBD Images
- Pedestrian Detection with Autoregressive Network Phases
- A Deep Multi-task Learning Approach to Skin Lesion Classification
- Detecting Pulsars with Neural Networks: A Proof of Concept
- Estimating Atmospheric Parameters of DA White Dwarf Stars with Deep Learning
- Table-Based Neural Units: Fully Quantizing Networks for Multiply-Free Inference
- Malaria detection from RBC images using shallow Convolutional Neural Networks
- Knowledge Distillation via Instance-level Sequence Learning
- Type-Driven Automated Learning with Lale
- Learning Visual Representations for Transfer Learning by Suppressing Texture
- Efficient Realistic Data Generation Framework leveraging Deep Learning-based Human Digitization
- A study on using image based machine learning methods to develop the surrogate models of stamp forming simulations
- VommaNet: an End-to-End Network for Disparity Estimation from Reflective and Texture-less Light Field Images
- Recognizing bird species in diverse soundscapes under weak supervision
- AttendSeg: A Tiny Attention Condenser Neural Network for Semantic Segmentation on the Edge
- CSMCNet: Scalable Video Compressive Sensing Reconstruction with Interpretable Motion Estimation
- Widget Captioning: Generating Natural Language Description for Mobile User Interface Elements
- FaceShapeGene: A Disentangled Shape Representation for Flexible Face Image Editing
- Guided Proofreading of Automatic Segmentations for Connectomics
- DeepGeo: Photo Localization with Deep Neural Network
- Spatial Memory for Context Reasoning in Object Detection
- Improving Sentence Representations with Consensus Maximisation
- Learning Disentangled Representations of Satellite Image Time Series
- A Deep Learning Approach to the Inversion of Borehole Resistivity Measurements
- iShape: A First Step Towards Irregular Shape Instance Segmentation
- Investigating Transfer Learning Capabilities of Vision Transformers and CNNs by Fine-Tuning a Single Trainable Block
- In Defense of the Classification Loss for Person Re-Identification
- MLPerf HPC: A Holistic Benchmark Suite for Scientific Machine Learning on HPC Systems
- Predicting the dynamics of 2d objects with a deep residual network
- Weighted Empirical Risk Minimization: Sample Selection Bias Correction based on Importance Sampling
- Representation Transfer by Optimal Transport
- Lossless Compression of Mosaic Images with Convolutional Neural Network Prediction
- Gradually Updated Neural Networks for Large-Scale Image Recognition
- An Audio-Video Deep and Transfer Learning Framework for Multimodal Emotion Recognition in the wild
- Towards Trainable Saliency Maps in Medical Imaging
- Ultrasound Diagnosis of COVID-19: Robustness and Explainability
- AON: Towards Arbitrarily-Oriented Text Recognition
- A Singular Value Perspective on Model Robustness
- Ignition: An End-to-End Supervised Model for Training Simulated Self-Driving Vehicles
- Benchmarking Inference Performance of Deep Learning Models on Analog Devices
- An Effective Hit-or-Miss Layer Favoring Feature Interpretation as Learned Prototypes Deformations
- Hierarchical Representation Network for Steganalysis of QIM Steganography in Low-Bit-Rate Speech Signals
- Cube Padding for Weakly-Supervised Saliency Prediction in 360° Videos
- Giving Commands to a Self-driving Car: A Multimodal Reasoner for Visual Grounding
- Enhanced Object Detection via Fusion With Prior Beliefs from Image Classification
- NWT: Towards natural audio-to-video generation with representation learning
- Multivariate-Information Adversarial Ensemble for Scalable Joint Distribution Matching
- ArchNet: Data Hiding Model in Distributed Machine Learning System
- Imaginative Walks: Generative Random Walk Deviation Loss for Improved Unseen Learning Representation
- SuperNet -- An efficient method of neural networks ensembling
- Deep CSI Learning for Gait Biometric Sensing and Recognition
- Towards Collaborative Intelligence Friendly Architectures for Deep Learning
- Classifying Tweet Sentiment Using the Hidden State and Attention Matrix of a Fine-tuned BERTweet Model
- Probabilistic Oriented Object Detection in Automotive Radar
- Physical Primitive Decomposition
- A 3D CNN Network with BERT For Automatic COVID-19 Diagnosis From CT-Scan Images
- Deep Frame Interpolation
- Optimizing Deep Neural Networks with Multiple Search Neuroevolution
- Learning image from projection: a full-automatic reconstruction (FAR) net for sparse-views computed tomography
- An Artificial Intelligence-Driven Agent for Real-Time Head-and-Neck IMRT Plan Generation using Conditional Generative Adversarial Network (cGAN)
- Click Here: Human-Localized Keypoints as Guidance for Viewpoint Estimation
- Meta-Learning Bidirectional Update Rules
- Evaluating Text-to-Image Matching using Binary Image Selection (BISON)
- Is a Picture Worth Ten Thousand Words in a Review Dataset?
- Multi-Plateau Ensemble for Endoscopic Artefact Segmentation and Detection
- Empirical Analysis of Image Caption Generation using Deep Learning
- An Overview on Data Representation Learning: From Traditional Feature Learning to Recent Deep Learning
- Defective samples simulation through Neural Style Transfer for automatic surface defect segment
- Generative-Discriminative Complementary Learning
- End-to-end learning potentials for structured attribute prediction
- Learning On-Road Visual Control for Self-Driving Vehicles with Auxiliary Tasks
- Sparsity in Deep Neural Networks - An Empirical Investigation with TensorQuant
- NavigationNet: A Large-scale Interactive Indoor Navigation Dataset
- cuConv: A CUDA Implementation of Convolution for CNN Inference
- Go with the Flow: Adaptive Control for Neural ODEs
- Large, fast and accurate HI intensity maps with latent overlap diffusion
- Fast and Accurate Online Video Object Segmentation via Tracking Parts
- Dendron: Enhancing Human Activity Recognition with On-Device TinyML Learning
- Scalable Metric Learning via Weighted Approximate Rank Component Analysis
- "What happens if..." Learning to Predict the Effect of Forces in Images
- Real-Time Cell Sorting with Scalable In Situ FPGA-Accelerated Deep Learning
- Class Subset Selection for Transfer Learning using Submodularity
- Incremental Learning in Person Re-Identification
- Dynamic Video Segmentation Network
- DeMeshNet: Blind Face Inpainting for Deep MeshFace Verification
- High-speed Railway Fastener Detection and Localization Method based on convolutional neural network
- Instance Scale Normalization for image understanding
- Geometrically Enriched Latent Spaces
- CrescendoNet: A Simple Deep Convolutional Neural Network with Ensemble Behavior
- Deep unsupervised domain adaptation applied to the Cherenkov Telescope Array Large-Sized Telescope
- Federated Learning over Wireless Networks: A Band-limited Coordinated Descent Approach
- Investigation of a Machine learning methodology for the SKA pulsar search pipeline
- Deep Imbalanced Attribute Classification using Visual Attention Aggregation
- MULE: Multimodal Universal Language Embedding
- Single Image Depth Estimation Trained via Depth from Defocus Cues
- Clothing Retrieval with Visual Attention Model
- Functional Error Correction for Robust Neural Networks
- Proxy Synthesis: Learning with Synthetic Classes for Deep Metric Learning
- Spatially-Adaptive Filter Units for Deep Neural Networks
- MLPerf Mobile Inference Benchmark
- Experimental Demonstration of Learned Time-Domain Digital Back-Propagation
- Deep Consensus Learning
- First-Order Preconditioning via Hypergradient Descent
- Single Sample Feature Importance: An Interpretable Algorithm for Low-Level Feature Analysis
- Hunt for dark subhalos in the galactic stellar field using computer vision
- BENCHIP: Benchmarking Intelligence Processors
- Where Should We Begin? A Low-Level Exploration of Weight Initialization Impact on Quantized Behaviour of Deep Neural Networks
- Looking for change? Roll the Dice and demand Attention
- Adaptive Leader-Follower Formation Control and Obstacle Avoidance via Deep Reinforcement Learning
- Class-Wise Difficulty-Balanced Loss for Solving Class-Imbalance
- Learning Spatial-Semantic Context with Fully Convolutional Recurrent Network for Online Handwritten Chinese Text Recognition
- Deep Transfer Learning with Ridge Regression
- On Improving Temporal Consistency for Online Face Liveness Detection
- Ensemble Transfer Learning for Emergency Landing Field Identification on Moderate Resource Heterogeneous Kubernetes Cluster
- MPDCompress - Matrix Permutation Decomposition Algorithm for Deep Neural Network Compression
- Quality-Aware Network for Human Parsing
- PoseGAN: A Pose-to-Image Translation Framework for Camera Localization
- Comparing the costs of abstraction for DL frameworks
- Recommending Outfits from Personal Closet
- Propagated Perturbation of Adversarial Attack for well-known CNNs: Empirical Study and its Explanation
- Learning View Priors for Single-view 3D Reconstruction
- Towards DeepSentinel: An extensible corpus of labelled Sentinel-1 and -2 imagery and a general-purpose sensor-fusion semantic embedding model
- Trying Bilinear Pooling in Video-QA
- NASIB: Neural Architecture Search withIn Budget
- El-CID: A filter for Gravitational-wave Electromagnetic Counterpart Identification
- Vehicle Re-identification Method Based on Vehicle Attribute and Mutual Exclusion Between Cameras
- Aug3D-RPN: Improving Monocular 3D Object Detection by Synthetic Images with Virtual Depth
- Energy Propagation in Deep Convolutional Neural Networks
- MinMaxCAM: Improving object coverage for CAM-basedWeakly Supervised Object Localization
- SAM-GCNN: A Gated Convolutional Neural Network with Segment-Level Attention Mechanism for Home Activity Monitoring
- Deconfusing intensity maps with neural networks
- Few-Shot Object Recognition from Machine-Labeled Web Images
- One-class Steel Detector Using Patch GAN Discriminator for Visualising Anomalous Feature Map
- Difficulty Translation in Histopathology Images
- YUVMultiNet: Real-time YUV multi-task CNN for autonomous driving
- Exploring Vision Transformers for Fine-grained Classification
- Meta-Cal: Well-controlled Post-hoc Calibration by Ranking
- Bespoke vs. Prêt-à-Porter Lottery Tickets: Exploiting Mask Similarity for Trainable Sub-Network Finding
- Real-time Burst Photo Selection Using a Light-Head Adversarial Network
- Explain Me the Painting: Multi-Topic Knowledgeable Art Description Generation
- Dense xUnit Networks
- WaveletNet: Logarithmic Scale Efficient Convolutional Neural Networks for Edge Devices
- Key Frame Proposal Network for Efficient Pose Estimation in Videos
- The VQA-Machine: Learning How to Use Existing Vision Algorithms to Answer New Questions
- Hybrid graph convolutional neural networks for landmark-based anatomical segmentation
- Model Optimization for Deep Space Exploration via Simulators and Deep Learning
- Understanding Catastrophic Forgetting and Remembering in Continual Learning with Optimal Relevance Mapping
- Grounded Video Description
- From Deep to Shallow: Transformations of Deep Rectifier Networks
- Compression Fractures Detection on CT
- Gear Training: A new way to implement high-performance model-parallel training
- A comparison of deep machine learning algorithms in COVID-19 disease diagnosis
- A Simple and Interpretable Predictive Model for Healthcare
- Stochastic Gradient Variance Reduction by Solving a Filtering Problem
- Improving compute efficacy frontiers with SliceOut
- Augmentation Inside the Network
- Reliable Label Bootstrapping for Semi-Supervised Learning
- A Machine-Learning Method for Time-Dependent Wave Equations over Unbounded Domains
- Feudal Steering: Hierarchical Learning for Steering Angle Prediction
- Adversarially Robust Kernel Smoothing
- Unsupervised Continual Learning Via Pseudo Labels
- Deep Learning: Our Miraculous Year 1990-1991
- D-OccNet: Detailed 3D Reconstruction Using Cross-Domain Learning
- DEPARA: Deep Attribution Graph for Deep Knowledge Transferability
- Improving Convolutional Neural Networks Via Conservative Field Regularisation and Integration
- Lenses In VoicE (LIVE): Searching for strong gravitational lenses in the VOICE@VST survey using Convolutional Neural Networks
- Multimodal Data Fusion based on the Global Workspace Theory
- C-DLinkNet: considering multi-level semantic features for human parsing
- Fine-grained Image-to-Image Transformation towards Visual Recognition
- ADNet: Leveraging Error-Bias Towards Normal Direction in Face Alignment
- What augmentations are sensitive to hyper-parameters and why?
- Cross-Channel Intragroup Sparsity Neural Network
- BPPSA: Scaling Back-propagation by Parallel Scan Algorithm
- Filter Bank Regularization of Convolutional Neural Networks
- Kilonova Light Curve Parameter Estimation Using Likelihood-Free Inference
- Latent Sensor Fusion: Multimedia Learning of Physiological Signals for Resource-Constrained Devices
- On Spectral Properties of Gradient-based Explanation Methods
- Whole Slide Multiple Instance Learning for Predicting Axillary Lymph Node Metastasis
- Sex Detection in the Early Stage of Fertilized Chicken Eggs via Image Recognition
- Structured 2D Representation of 3D Data for Shape Processing
- Lift-the-flap: what, where and when for context reasoning
- Using Early-Learning Regularization to Classify Real-World Noisy Data
- Regularized Ensembles and Transferability in Adversarial Learning
- Reducing the feature divergence of RGB and near-infrared images using Switchable Normalization
- Neural Rejuvenation: Improving Deep Network Training by Enhancing Computational Resource Utilization
- D2C: Diffusion-Denoising Models for Few-shot Conditional Generation
- Orthogonal-Padé Activation Functions: Trainable Activation functions for smooth and faster convergence in deep networks
- Augmented KRnet for density estimation and approximation
- ChainGAN: A sequential approach to GANs
- Semantic Image Cropping
- Will Multi-modal Data Improves Few-shot Learning?
- Scene Parsing via Dense Recurrent Neural Networks with Attentional Selection
- Retinal Microvasculature as Biomarker for Diabetes and Cardiovascular Diseases
- Deep Learning At Scale and At Ease
- Semantic Segmentation and Object Detection Towards Instance Segmentation: Breast Tumor Identification
- The Method of Multimodal MRI Brain Image Segmentation Based on Differential Geometric Features
- Machine Translation between Vietnamese and English: an Empirical Study
- Effectiveness of Deep Networks in NLP using BiDAF as an example architecture
- Semantic Segmentation on VSPW Dataset through Aggregation of Transformer Models
- Learnable Pooling Methods for Video Classification
- Multiple Myeloma Cancer Cell Instance Segmentation
- Dense neural networks as sparse graphs and the lightning initialization
- VTAMIQ: Transformers for Attention Modulated Image Quality Assessment
- Disturbing Target Values for Neural Network Regularization
- Evaluation of Neural Networks for Image Recognition Applications: Designing a 0-1 MILP Model of a CNN to create adversarials
- Protecting Anonymous Speech: A Generative Adversarial Network Methodology for Removing Stylistic Indicators in Text
- Exploiting Inter-pixel Correlations in Unsupervised Domain Adaptation for Semantic Segmentation
- Smart Device based Initial Movement Detection of Cyclists using Convolutional Neuronal Networks
- LogAvgExp Provides a Principled and Performant Global Pooling Operator
- Auto-Classification of Retinal Diseases in the Limit of Sparse Data Using a Two-Streams Machine Learning Model
- Fully convolutional Siamese neural networks for buildings damage assessment from satellite images
- Unifying Identification and Context Learning for Person Recognition
- Convolution in Convolution for Network in Network
- Foresee: Attentive Future Projections of Chaotic Road Environments with Online Training
- Doping: A technique for efficient compression of LSTM models using sparse structured additive matrices
- The Effects of Image Distribution and Task on Adversarial Robustness
- A Neural Embeddings Approach for Detecting Mobile Counterfeit Apps
- A Sketch-Based Neural Model for Generating Commit Messages from Diffs
- Writing in The Air: Unconstrained Text Recognition from Finger Movement Using Spatio-Temporal Convolution
- Automatic Latent Fingerprint Segmentation
- Deep Collaborative Learning for Visual Recognition
- Decoupled Gradient Harmonized Detector for Partial Annotation: Application to Signet Ring Cell Detection
- Joint Spatial-Temporal Optimization for Stereo 3D Object Tracking
- Intensity Non-uniformity Correction in MR Imaging Using Residual Cycle Generative Adversarial Network
- In-Vehicle Object Detection in the Wild for Driverless Vehicles
- Model Agnostic Combination for Ensemble Learning
- Amended Cross Entropy Cost: Framework For Explicit Diversity Encouragement
- Non-Markov Policies to Reduce Sequential Failures in Robot Bin Picking
- Unsupervised Learning of Solutions to Differential Equations with Generative Adversarial Networks
- Deep Learning with Apache SystemML
- Detecting abnormal events in video using Narrowed Normality Clusters
- A general approach to bridge the reality-gap
- Enforcing Reasoning in Visual Commonsense Reasoning
- Towards thinner convolutional neural networks through Gradually Global Pruning
- TraffickCam: Explainable Image Matching For Sex Trafficking Investigations
- Incorporating Textual Evidence in Visual Storytelling
- Privacy for Rescue: A New Testimony Why Privacy is Vulnerable In Deep Models
- Adaptive Loss Function for Super Resolution Neural Networks Using Convex Optimization Techniques
- NeuroFabric: Identifying Ideal Topologies for Training A Priori Sparse Networks
- Adversarial training applied to Convolutional Neural Network for photometric redshift predictions
- Deep Shape Matching
- FA-RPN: Floating Region Proposals for Face Detection
- Studying the Plasticity in Deep Convolutional Neural Networks using Random Pruning
- Skeleton Transformer Networks: 3D Human Pose and Skinned Mesh from Single RGB Image
- Scalable and Effective Deep CCA via Soft Decorrelation
- Learning scale-variant and scale-invariant features for deep image classification
- A Push-Pull Layer Improves Robustness of Convolutional Neural Networks
- A Generative Map for Image-based Camera Localization
- Manifestation of Image Contrast in Deep Networks
- Human Activity Recognition for Edge Devices
- Boosted Attention: Leveraging Human Attention for Image Captioning
- Exploring epoch-dependent stochastic residual networks
- Deep Contextual Recurrent Residual Networks for Scene Labeling
- Perceive Where to Focus: Learning Visibility-aware Part-level Features for Partial Person Re-identification
- Position-Aware Convolutional Networks for Traffic Prediction
- Unsupervised Open Domain Recognition by Semantic Discrepancy Minimization
- ClusterNet: Detecting Small Objects in Large Scenes by Exploiting Spatio-Temporal Information
- Beyond Planar Symmetry: Modeling human perception of reflection and rotation symmetries in the wild
- Positional Artefacts Propagate Through Masked Language Model Embeddings
- YouTube-8M Video Understanding Challenge Approach and Applications
- Recurrent 3D Pose Sequence Machines
- High Efficient Reconstruction of Single-shot T2 Mapping from OverLapping-Echo Detachment Planar Imaging Based on Deep Residual Network
- BPGrad: Towards Global Optimality in Deep Learning via Branch and Pruning
- Feature Selective Networks for Object Detection
- Multi-Mode Inference Engine for Convolutional Neural Networks
- Visual Data Augmentation through Learning
- Deep, Dense, and Low-Rank Gaussian Conditional Random Fields
- TLGAN: document Text Localization using Generative Adversarial Nets
- Audiomer: A Convolutional Transformer For Keyword Spotting
- A Simple Riemannian Manifold Network for Image Set Classification
- Autoencoding Undirected Molecular Graphs With Neural Networks
- LiDAR-based Recurrent 3D Semantic Segmentation with Temporal Memory Alignment
- BEDS: Bagging ensemble deep segmentation for nucleus segmentation with testing stage stain augmentation
- Population Gradients improve performance across data-sets and architectures in object classification
- Hardware-efficient Residual Networks for FPGAs
- Rethinking Radiology: An Analysis of Different Approaches to BraTS
- VideoClick: Video Object Segmentation with a Single Click
- RethNet: Object-by-Object Learning for Detecting Facial Skin Problems
- Bayesian neural network with pretrained protein embedding enhances prediction accuracy of drug-protein interaction
- Diagnosis of Pediatric Obstructive Sleep Apnea via Face Classification with Persistent Homology and Convolutional Neural Networks
- A Multi-Task Learning Approach for Meal Assessment
- MyFood: A Food Segmentation and Classification System to Aid Nutritional Monitoring
- Selective Deep Convolutional Neural Network for Low Cost Distorted Image Classification
- Leveraging Model Interpretability and Stability to increase Model Robustness
- Investigations of the Influences of a CNN's Receptive Field on Segmentation of Subnuclei of Bilateral Amygdalae
- Vehicle Image Generation Going Well with The Surroundings
- Dense Fusion Classmate Network for Land Cover Classification
- Fast and Explicit Neural View Synthesis
- 3D attention mechanism for fine-grained classification of table tennis strokes using a Twin Spatio-Temporal Convolutional Neural Networks
- FusionStitching: Boosting Execution Efficiency of Memory Intensive Computations for DL Workloads
- Medusa: A Scalable Interconnect for Many-Port DNN Accelerators and Wide DRAM Controller Interfaces
- Learning to Predict the 3D Layout of a Scene
- Granular Motor State Monitoring of Free Living Parkinson's Disease Patients via Deep Learning
- Relational Mimic for Visual Adversarial Imitation Learning
- Bounding Box Embedding for Single Shot Person Instance Segmentation
- RootPainter3D: Interactive-machine-learning enables rapid and accurate contouring for radiotherapy
- Text Classification based on Multiple Block Convolutional Highways
- Temporally Folded Convolutional Neural Networks for Sequence Forecasting
- On the Flip Side: Identifying Counterexamples in Visual Question Answering
- Person re-identification across different datasets with multi-task learning
- Towards More Efficient and Effective Inference: The Joint Decision of Multi-Participants
- Person Search in Videos with One Portrait Through Visual and Temporal Links
- Learning Identity-Preserving Transformations on Data Manifolds
- Mitigating large adversarial perturbations on X-MAS (X minus Moving Averaged Samples)
- Phantom: A High-Performance Computational Core for Sparse Convolutional Neural Networks
- Recent Advancements in Self-Supervised Paradigms for Visual Feature Representation
- A Survey On 3D Inner Structure Prediction from its Outer Shape
- OpTorch: Optimized deep learning architectures for resource limited environments
- Backtracking gradient descent method for general functions, with applications to Deep Learning
- A simple model for detection of rare sound events
- Deep Boosted Regression for MR to CT Synthesis
- A one-armed CNN for exoplanet detection from light curves
- ReflectNet -- A Generative Adversarial Method for Single Image Reflection Suppression
- FMT:Fusing Multi-task Convolutional Neural Network for Person Search
- Video Surveillance for Road Traffic Monitoring
- Attentive Sequence to Sequence Translation for Localizing Clips of Interest by Natural Language Descriptions
- MCMC Guided CNN Training and Segmentation for Pancreas Extraction
- A Content Transformation Block For Image Style Transfer
- Unsupervised monocular stereo matching
- 3D-MOV: Audio-Visual LSTM Autoencoder for 3D Reconstruction of Multiple Objects from Video
- Image Manipulation with Perceptual Discriminators
- Deep Generative Variational Autoencoding for Replay Spoof Detection in Automatic Speaker Verification
- Exploration on Grounded Word Embedding: Matching Words and Images with Image-Enhanced Skip-Gram Model
- Bidirectional Multirate Reconstruction for Temporal Modeling in Videos
- Hotels-50K: A Global Hotel Recognition Dataset
- Open Source Face Recognition Performance Evaluation Package
- TUNet: Incorporating segmentation maps to improve classification
- Dual Pattern Learning Networks by Empirical Dual Prediction Risk Minimization
- LR-to-HR Face Hallucination with an Adversarial Progressive Attribute-Induced Network
- Barrier-Free Large-Scale Sparse Tensor Accelerator (BARISTA) For Convolutional Neural Networks
- Analysing object detectors from the perspective of co-occurring object categories
- Classification of COVID-19 from CXR Images in a 15-class Scenario: an Attempt to Avoid Bias in the System
- Improving the Performance of Neural Networks in Regression Tasks Using Drawering
- Animal inspired Application of a Variant of Mel Spectrogram for Seismic Data Processing
- Reconciling Feature-Reuse and Overfitting in DenseNet with Specialized Dropout
- Deep Active Localization
- Benchmarks of ResNet Architecture for Atrial Fibrillation Classification
- Asynchronous Stochastic Gradient MCMC with Elastic Coupling
- Accent Recognition with Hybrid Phonetic Features
- Introduce the Result Into Self-Attention
- Automated pulmonary nodule detection using 3D deep convolutional neural networks
- Perceptual Gradient Networks
- Training Dynamic based data filtering may not work for NLP datasets
- Finding Correspondences for Optical Flow and Disparity Estimations using a Sub-pixel Convolution-based Encoder-Decoder Network
- Natural Disaster Classification using Aerial Photography Explainable for Typhoon Damaged Feature
- Minimizing Labeling Effort for Tree Skeleton Segmentation using an Automated Iterative Training Methodology
- Training Convolutional Neural Networks and Compressed Sensing End-to-End for Microscopy Cell Detection
- Penetrating the Fog: the Path to Efficient CNN Models
- The Focus-Aspect-Polarity Model for Predicting Subjective Noun Attributes in Images
- From Artificial Intelligence to Brain Intelligence: The basis learning and memory algorithm for brain-like intelligence
- StackGAN: Facial Image Generation Optimizations
- GANs for Urban Design
- Real-time Action Recognition with Dissimilarity-based Training of Specialized Module Networks
- On the Equivalence of Convolutional and Hadamard Networks using DFT
- Structure Learning of Deep Neural Networks with Q-Learning
- TrUMAn: Trope Understanding in Movies and Animations
- Logit Attenuating Weight Normalization
- Copy and Paste method based on Pose for Re-identification
- Development and evaluation of intraoperative ultrasound segmentation with negative image frames and multiple observer labels
- Improve Unsupervised Pretraining for Few-label Transfer
- Style-Restricted GAN: Multi-Modal Translation with Style Restriction Using Generative Adversarial Networks
- Learning Robust 3D Face Reconstruction and Discriminative Identity Representation
- 2nd Place Solution to Instance Segmentation of IJCAI 3D AI Challenge 2020
- Batch Inverse-Variance Weighting: Deep Heteroscedastic Regression
- Mix and Mask Actor-Critic Methods
- Y-GAN: A Generative Adversarial Network for Depthmap Estimation from Multi-camera Stereo Images
- A Compositional Textual Model for Recognition of Imperfect Word Images
- Automatic Liver Segmentation with Adversarial Loss and Convolutional Neural Network
- Pay Attention to Convolution Filters: Towards Fast and Accurate Fine-Grained Transfer Learning
- Binary Stochastic Representations for Large Multi-class Classification
- Semi-Supervised Learning for Face Sketch Synthesis in the Wild
- Playing Go without Game Tree Search Using Convolutional Neural Networks
- Neural Image Captioning
- Vanishing Curvature and the Power of Adaptive Methods in Randomly Initialized Deep Networks
- DeepCompress: Efficient Point Cloud Geometry Compression
- MURAUER: Mapping Unlabeled Real Data for Label AUstERity
- DeepAtrophy: Teaching a Neural Network to Differentiate Progressive Changes from Noise on Longitudinal MRI in Alzheimer's Disease
- Montage based 3D Medical Image Retrieval from Traumatic Brain Injury Cohort using Deep Convolutional Neural Network
- Diagnostic Visualization for Deep Neural Networks Using Stochastic Gradient Langevin Dynamics
- Know Your Surroundings: Panoramic Multi-Object Tracking by Multimodality Collaboration
- Does Face Recognition Error Echo Gender Classification Error?
- Neural Machine Translation between Herbal Prescriptions and Diseases
- Towards Diverse Paragraph Captioning for Untrimmed Videos
- Neural Options Pricing
- Efficient Transfer Learning via Joint Adaptation of Network Architecture and Weight
- Mapping the Internet: Modelling Entity Interactions in Complex Heterogeneous Networks
- Domain-Agnostic Clustering with Self-Distillation
- Towards Robust and Automatic Hyper-Parameter Tunning
- An Improved Neural Segmentation Method Based on U-NET
- Learning Chebyshev Basis in Graph Convolutional Networks for Skeleton-based Action Recognition
- Large Language Models -- the Future of Fundamental Physics?
- InsectUp: Crowdsourcing Insect Observations to Assess Demographic Shifts and Improve Classification
- Tensorization of neural networks for improved privacy and interpretability
- Searches for Compact Binary Coalescence Events Using Neural Networks in LIGO/Virgo Third Observation Period
- Image to Video Domain Adaptation Using Web Supervision
- DeepWheat: Estimating Phenotypic Traits from Crop Images with Deep Learning
- Variational Autoencoded Regression: High Dimensional Regression of Visual Data on Complex Manifold
- Non-local RoIs for Instance Segmentation
- Real-time Human Detection Model for Edge Devices
- A deep ensemble approach to X-ray polarimetry
- Using mixup as regularization and tuning hyper-parameters for ResNets
- Recognition of Russian traffic signs in winter conditions. Solutions of the "Ice Vision" competition winners
- Depth Without the Magic: Inductive Bias of Natural Gradient Descent
- Consistent Accelerated Inference via Confident Adaptive Transformers
- Block-Cyclic Stochastic Coordinate Descent for Deep Neural Networks
- Who Will Win Practical Artificial Intelligence? AI Engineerings in China
- Train, Diagnose and Fix: Interpretable Approach for Fine-grained Action Recognition
- Predicting Clinical Outcomes in COVID-19 using Radiomics and Deep Learning on Chest Radiographs: A Multi-Institutional Study
- Hybrid BYOL-ViT: Efficient approach to deal with small datasets
- Exploiting Nontrivial Connectivity for Automatic Speech Recognition
- Evaluating Contrastive Learning on Wearable Timeseries for Downstream Clinical Outcomes
- Learning by Cheating : An End-to-End Zero Shot Framework for Autonomous Drone Navigation
- Graph Relation Transformer: Incorporating pairwise object features into the Transformer architecture
- Complementary Ensemble Learning
- DDNet: Dual-path Decoder Network for Occlusion Relationship Reasoning
- Game Theory for Adversarial Attacks and Defenses
- A Self-supervised Learning System for Object Detection in Videos Using Random Walks on Graphs
- FHEDN: A based on context modeling Feature Hierarchy Encoder-Decoder Network for face detection
- A Free Lunch in Generating Datasets: Building a VQG and VQA System with Attention and Humans in the Loop
- Split-Merge Pooling
- Searching Learning Strategy with Reinforcement Learning for 3D Medical Image Segmentation
- Knowledge-driven Active Learning
- Sparsifying and Down-scaling Networks to Increase Robustness to Distortions
- Recurrent Residual Module for Fast Inference in Videos
- Error Control and Loss Functions for the Deep Learning Inversion of Borehole Resistivity Measurements
- Exponential Discriminative Metric Embedding in Deep Learning
- On Intrinsic Dataset Properties for Adversarial Machine Learning
- Cross-filter compression for CNN inference acceleration
- PrototypeML: A Neural Network Integrated Design and Development Environment
- A Multi-Modal Approach to Infer Image Affect
- Efficient Embedding of MPI Collectives in MXNET DAGs for scaling Deep Learning
- Learning Atomic Multipoles: Prediction of the Electrostatic Potential with Equivariant Graph Neural Networks
- Train and Deploy an Image Classifier for Disaster Response
- Survey on Deep Learning-based Kuzushiji Recognition
- Broadband vectorial ultra-flat optics with experimental efficiency up to 99% in the visible via universal approximators
- Identify Speakers in Cocktail Parties with End-to-End Attention
- Rethinking Lightweight Convolutional Neural Networks for Efficient and High-quality Pavement Crack Detection
- ISyNet: Convolutional Neural Networks design for AI accelerator
- Local Contrast Learning
- Caramel: Accelerating Decentralized Distributed Deep Learning with Computation Scheduling
- Compact retail shelf segmentation for mobile deployment
- Impact of Channel Variation on One-Class Learning for Spoof Detection
- How to Train your DNN: The Network Operator Edition
- VeriMedi: Pill Identification using Proxy-based Deep Metric Learning and Exact Solution
- OmniLayout: Room Layout Reconstruction from Indoor Spherical Panoramas
- Automated Segmentation of Brain Gray Matter Nuclei on Quantitative Susceptibility Mapping Using Deep Convolutional Neural Network
- Holistic Image Manipulation Detection using Pixel Co-occurrence Matrices
- Efficient resource management in UAVs for Visual Assistance
- Domain Adaptive Monocular Depth Estimation With Semantic Information
- Reinforced Attention for Few-Shot Learning and Beyond
- Representation range needs for 16-bit neural network training
- Convergence Analysis of Gradient Descent Algorithms with Proportional Updates
- Graph Neural Networks for UnsupervisedDomain Adaptation of Histopathological ImageAnalytics
- Beyond Categorical Label Representations for Image Classification
- Exploiting Temporal Attention Features for Effective Denoising in Videos
- Analysis on Image Set Visual Question Answering
- Learning Compact Physics-Aware Delayed Photocurrent Models Using Dynamic Mode Decomposition
- AET-EFN: A Versatile Design for Static and Dynamic Event-Based Vision
- Instance Map based Image Synthesis with a Denoising Generative Adversarial Network
- Dynamic Metric Learning: Towards a Scalable Metric Space to Accommodate Multiple Semantic Scales
- A Quadratic Actor Network for Model-Free Reinforcement Learning
- Deep Gradient Projection Networks for Pan-sharpening