Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
arXiv:1506.01497
Abstract
State-of-the-art object detection networks depend on region proposal algorithms to hypothesize object locations. Advances like SPPnet and Fast R-CNN have reduced the running time of these detection networks, exposing region proposal computation as a bottleneck. In this work, we introduce a Region Proposal Network (RPN) that shares full-image convolutional features with the detection network, thus enabling nearly cost-free region proposals. An RPN is a fully convolutional network that simultaneously predicts object bounds and objectness scores at each position. The RPN is trained end-to-end to generate high-quality region proposals, which are used by Fast R-CNN for detection. We further merge RPN and Fast R-CNN into a single network by sharing their convolutional features---using the recently popular terminology of neural networks with 'attention' mechanisms, the RPN component tells the unified network where to look. For the very deep VGG-16 model, our detection system has a frame rate of 5fps (including all steps) on a GPU, while achieving state-of-the-art object detection accuracy on PASCAL VOC 2007, 2012, and MS COCO datasets with only 300 proposals per image. In ILSVRC and COCO 2015 competitions, Faster R-CNN and RPN are the foundations of the 1st-place winning entries in several tracks. Code has been made publicly available.
Extended tech report
References in corpus (4)
Cited by in corpus (3083)
- YOLOv4: Optimal Speed and Accuracy of Object Detection
- Deep Residual Learning for Image Recognition
- A Survey on Visual Transformer
- Simple Online and Realtime Tracking
- Transformers in Vision: A Survey
- Res2Net: A New Multi-scale Backbone Architecture
- Bootstrap your own latent: A new approach to self-supervised Learning
- YOLOX: Exceeding YOLO Series in 2021
- MobileNetV2: Inverted Residuals and Linear Bottlenecks
- Improved Baselines with Momentum Contrastive Learning
- HybridSN: Exploring 3D-2D CNN Feature Hierarchy for Hyperspectral Image Classification
- Deformable DETR: Deformable Transformers for End-to-End Object Detection
- GOT-10k: A Large High-Diversity Benchmark for Generic Object Tracking in the Wild
- UnitBox: An Advanced Object Detection Network
- FairMOT: On the Fairness of Detection and Re-Identification in Multiple Object Tracking
- ADVENT: Adversarial Entropy Minimization for Domain Adaptation in Semantic Segmentation
- Semantic Foggy Scene Understanding with Synthetic Data
- AMC: AutoML for Model Compression and Acceleration on Mobile Devices
- VisualBERT: A Simple and Performant Baseline for Vision and Language
- Person Re-identification: Past, Present and Future
- Momentum Contrast for Unsupervised Visual Representation Learning
- The VIA Annotation Software for Images, Audio and Video
- Distance-IoU Loss: Faster and Better Learning for Bounding Box Regression
- All One Needs to Know about Metaverse: A Complete Survey on Technological Singularity, Virtual Ecosystem, and Research Agenda
- Contrastive Representation Learning: A Framework and Review
- FCOS: Fully Convolutional One-Stage Object Detection
- Multi-Task Learning for Dense Prediction Tasks: A Survey
- MMDetection: Open MMLab Detection Toolbox and Benchmark
- Richer Convolutional Features for Edge Detection
- Barlow Twins: Self-Supervised Learning via Redundancy Reduction
- Anabranch Network for Camouflaged Object Segmentation
- A Survey of Human-in-the-loop for Machine Learning
- DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
- Learning Representations by Maximizing Mutual Information Across Views
- OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields
- Video Salient Object Detection via Fully Convolutional Networks
- What Makes for Good Views for Contrastive Learning?
- The ApolloScape Open Dataset for Autonomous Driving and its Application
- CutMix: Regularization Strategy to Train Strong Classifiers with Localizable Features
- T-CNN: Tubelets with Convolutional Neural Networks for Object Detection from Videos
- DeepSaliency: Multi-Task Deep Neural Network Model for Salient Object Detection
- Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges
- This Looks Like That: Deep Learning for Interpretable Image Recognition
- Binary Neural Networks: A Survey
- YOLACT++: Better Real-time Instance Segmentation
- EfficientDet: Scalable and Efficient Object Detection
- Evaluate the Malignancy of Pulmonary Nodules Using the 3D Deep Leaky Noisy-or Network
- NWPU-Crowd: A Large-Scale Benchmark for Crowd Counting and Localization
- FPGA-based Accelerators of Deep Learning Networks for Learning and Classification: A Review
- Prototypical Contrastive Learning of Unsupervised Representations
- Deep Clustering for Unsupervised Learning of Visual Features
- 3D Object Detection for Autonomous Driving: A Survey
- YOLOP: You Only Look Once for Panoptic Driving Perception
- Image-based 3D Object Reconstruction: State-of-the-Art and Trends in the Deep Learning Era
- Deep Continuous Fusion for Multi-Sensor 3D Object Detection
- The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale
- RetinaFace: Single-stage Dense Face Localisation in the Wild
- Detecting tiny objects in aerial images: A normalized Wasserstein distance and a new benchmark
- Exploring Simple Siamese Representation Learning
- Facial Landmark Detection: a Literature Survey
- Real-Time Polyp Detection, Localization and Segmentation in Colonoscopy Using Deep Learning
- Aligning Domain-specific Distribution and Classifier for Cross-domain Classification from Multiple Sources
- Aggregated Residual Transformations for Deep Neural Networks
- Deep Learning for UAV-based Object Detection and Tracking: A Survey
- TransTrack: Multiple Object Tracking with Transformer
- SINet: A Scale-insensitive Convolutional Neural Network for Fast Vehicle Detection
- INTERACTION Dataset: An INTERnational, Adversarial and Cooperative moTION Dataset in Interactive Driving Scenarios with Semantic Maps
- Early Convolutions Help Transformers See Better
- Image Classification with Deep Learning in the Presence of Noisy Labels: A Survey
- Fast and Furious: Real Time End-to-End 3D Detection, Tracking and Motion Forecasting with a Single Convolutional Net
- SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
- Cooperative Perception for 3D Object Detection in Driving Scenarios using Infrastructure Sensors
- A Simple Semi-Supervised Learning Framework for Object Detection
- A Review and Comparative Study on Probabilistic Object Detection in Autonomous Driving
- Multimodal End-to-End Autonomous Driving
- WebVision Database: Visual Learning and Understanding from Web Data
- The EuroCity Persons Dataset: A Novel Benchmark for Object Detection
- An Iterative BP-CNN Architecture for Channel Decoding
- LCR-Net++: Multi-person 2D and 3D Pose Detection in Natural Images
- Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
- Scaling Egocentric Vision: The EPIC-KITCHENS Dataset
- Working Memory Connections for LSTM
- VLP: A Survey on Vision-Language Pre-training
- TransCrowd: weakly-supervised crowd counting with transformers
- Hard Negative Mixing for Contrastive Learning
- Apple Flower Detection using Deep Convolutional Networks
- A Normalized Gaussian Wasserstein Distance for Tiny Object Detection
- Predicting Head Movement in Panoramic Video: A Deep Reinforcement Learning Approach
- Learning Object Bounding Boxes for 3D Instance Segmentation on Point Clouds
- Visual Question Answering: Datasets, Algorithms, and Future Challenges
- A Dataset And Benchmark Of Underwater Object Detection For Robot Picking
- BlazeFace: Sub-millisecond Neural Face Detection on Mobile GPUs
- Image De-raining Using a Conditional Generative Adversarial Network
- CFC-Net: A Critical Feature Capturing Network for Arbitrary-Oriented Object Detection in Remote Sensing Images
- Monitoring COVID-19 social distancing with person detection and tracking via fine-tuned YOLO v3 and Deepsort techniques
- Remote Sensing Image Super-resolution and Object Detection: Benchmark and State of the Art
- A Survey on Deep Learning-based Architectures for Semantic Segmentation on 2D images
- Learning High-Precision Bounding Box for Rotated Object Detection via Kullback-Leibler Divergence
- Deep Learning for Generic Object Detection: A Survey
- PP-YOLO: An Effective and Efficient Implementation of Object Detector
- GridMask Data Augmentation
- Weakly Supervised Adversarial Domain Adaptation for Semantic Segmentation in Urban Scenes
- RODNet: A Real-Time Radar Object Detection Network Cross-Supervised by Camera-Radar Fused Object 3D Localization
- LXMERT: Learning Cross-Modality Encoder Representations from Transformers
- Automatic Colon Polyp Detection using Region based Deep CNN and Post Learning Approaches
- Joint Activity Recognition and Indoor Localization with WiFi Fingerprints
- Stand-Alone Self-Attention in Vision Models
- Strengths and Weaknesses of Deep Learning Models for Face Recognition Against Image Degradations
- GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond
- Learning JPEG Compression Artifacts for Image Manipulation Detection and Localization
- Revisiting Anchor Mechanisms for Temporal Action Localization
- Real-Time Apple Detection System Using Embedded Systems With Hardware Accelerators: An Edge AI Application
- DeepIM: Deep Iterative Matching for 6D Pose Estimation
- BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning
- Driver Drowsiness Detection Model Using Convolutional Neural Networks Techniques for Android Application
- A Comprehensive Survey of Neural Architecture Search: Challenges and Solutions
- Weakly Supervised Fine-Grained Image Categorization
- Learning to Compose and Reason with Language Tree Structures for Visual Grounding
- Patch-VQ: 'Patching Up' the Video Quality Problem
- K-Net: Towards Unified Image Segmentation
- GhostNets on Heterogeneous Devices via Cheap Operations
- AutoAssign: Differentiable Label Assignment for Dense Object Detection
- Recurrently Exploring Class-wise Attention in A Hybrid Convolutional and Bidirectional LSTM Network for Multi-label Aerial Image Classification
- SCALE-Sim: Systolic CNN Accelerator Simulator
- Deep Hough Transform for Semantic Line Detection
- Deep Learning Techniques for Future Intelligent Cross-Media Retrieval
- Unifying Vision-and-Language Tasks via Text Generation
- An Attention-Fused Network for Semantic Segmentation of Very-High-Resolution Remote Sensing Imagery
- Gland Instance Segmentation Using Deep Multichannel Neural Networks
- Adversarial Example Detection for DNN Models: A Review and Experimental Comparison
- Spatial Group-wise Enhance: Improving Semantic Feature Learning in Convolutional Networks
- Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey
- Monocular Semantic Occupancy Grid Mapping with Convolutional Variational Encoder-Decoder Networks
- Visual Identification of Individual Holstein-Friesian Cattle via Deep Metric Learning
- iSAID: A Large-scale Dataset for Instance Segmentation in Aerial Images
- Hybrid Task Cascade for Instance Segmentation
- Deep learning for detection and segmentation of artefact and disease instances in gastrointestinal endoscopy
- Position Detection and Direction Prediction for Arbitrary-Oriented Ships via Multitask Rotation Region Convolutional Neural Network
- Deep-Learning-Based Image Segmentation Integrated with Optical Microscopy for Automatically Searching for Two-Dimensional Materials
- RGB-D Object Detection and Semantic Segmentation for Autonomous Manipulation in Clutter
- Threat of Adversarial Attacks on Deep Learning in Computer Vision: A Survey
- TSM: Temporal Shift Module for Efficient Video Understanding
- Simultaneous Traffic Sign Detection and Boundary Estimation using Convolutional Neural Network
- DeeperLab: Single-Shot Image Parser
- CornerNet-Lite: Efficient Keypoint Based Object Detection
- Universal Domain Adaptation through Self Supervision
- Probabilistic two-stage detection
- From Handcrafted to Deep Features for Pedestrian Detection: A Survey
- Efficient DETR: Improving End-to-End Object Detector with Dense Prior
- Dynamic Feature Integration for Simultaneous Detection of Salient Object, Edge and Skeleton
- From Google Maps to a Fine-Grained Catalog of Street trees
- STEm-Seg: Spatio-temporal Embeddings for Instance Segmentation in Videos
- A Survey on Long-Tailed Visual Recognition
- Graph-based Spatial-temporal Feature Learning for Neuromorphic Vision Sensing
- Context-Aware Visual Policy Network for Fine-Grained Image Captioning
- Deep Learning in Photoacoustic Tomography: Current approaches and future directions
- DetectoRS: Detecting Objects with Recursive Feature Pyramid and Switchable Atrous Convolution
- Audiovisual SlowFast Networks for Video Recognition
- BoxCars: Improving Fine-Grained Recognition of Vehicles using 3-D Bounding Boxes in Traffic Surveillance
- SegICP: Integrated Deep Semantic Segmentation and Pose Estimation
- Robust Spatial Filtering with Graph Convolutional Neural Networks
- Seeing Around Street Corners: Non-Line-of-Sight Detection and Tracking In-the-Wild Using Doppler Radar
- DPatch: An Adversarial Patch Attack on Object Detectors
- Multiscale Domain Adaptive YOLO for Cross-Domain Object Detection
- Bag of Freebies for Training Object Detection Neural Networks
- CornerNet: Detecting Objects as Paired Keypoints
- RepPoints: Point Set Representation for Object Detection
- Cascaded Region-based Densely Connected Network for Event Detection: A Seismic Application
- SlowFast Networks for Video Recognition
- Half a Percent of Labels is Enough: Efficient Animal Detection in UAV Imagery using Deep CNNs and Active Learning
- A CNN Approach to Simultaneously Count Plants and Detect Plantation-Rows from UAV Imagery
- CNN-based Density Estimation and Crowd Counting: A Survey
- UA-DETRAC: A New Benchmark and Protocol for Multi-Object Detection and Tracking
- Toward Transformer-Based Object Detection
- A deep learning framework for quality assessment and restoration in video endoscopy
- Convolutional Oriented Boundaries
- MirrorNet: Bio-Inspired Camouflaged Object Segmentation
- Real-Time Fruit Recognition and Grasping Estimation for Autonomous Apple Harvesting
- Libra R-CNN: Towards Balanced Learning for Object Detection
- An All-in-One Network for Dehazing and Beyond
- Joint 3D Proposal Generation and Object Detection from View Aggregation
- Towards Accurate One-Stage Object Detection with AP-Loss
- Masked-attention Mask Transformer for Universal Image Segmentation
- Benchmarking Robustness in Object Detection: Autonomous Driving when Winter is Coming
- Track to Reconstruct and Reconstruct to Track
- JRDB: A Dataset and Benchmark of Egocentric Robot Visual Perception of Humans in Built Environments
- Exploring Spatial Significance via Hybrid Pyramidal Graph Network for Vehicle Re-identification
- Learning Dual Semantic Relations with Graph Attention for Image-Text Matching
- FastDeepIoT: Towards Understanding and Optimizing Neural Network Execution Time on Mobile and Embedded Devices
- Learning to detect chest radiographs containing lung nodules using visual attention networks
- Counting of Grapevine Berries in Images via Semantic Segmentation using Convolutional Neural Networks
- Scaled-YOLOv4: Scaling Cross Stage Partial Network
- ICNet for Real-Time Semantic Segmentation on High-Resolution Images
- Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling
- Deep Learning for Logo Recognition
- Pix2seq: A Language Modeling Framework for Object Detection
- Pseudo-LiDAR++: Accurate Depth for 3D Object Detection in Autonomous Driving
- Drivers Drowsiness Detection using Condition-Adaptive Representation Learning Framework
- Incorporating User Micro-behaviors and Item Knowledge into Multi-task Learning for Session-based Recommendation
- Neural Networks for Entity Matching: A Survey
- TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios
- Learning Multilayer Channel Features for Pedestrian Detection
- Micro-Batch Training with Batch-Channel Normalization and Weight Standardization
- A Survey on ML4VIS: Applying Machine Learning Advances to Data Visualization
- Accelerating 3D Deep Learning with PyTorch3D
- Simple Online and Realtime Tracking with a Deep Association Metric
- Postdisaster image-based damage detection and repair cost estimation of reinforced concrete buildings using dual convolutional neural networks
- Rain rendering for evaluating and improving robustness to bad weather
- RetinaMask: Learning to predict masks improves state-of-the-art single-shot detection for free
- Image coding for machines: an end-to-end learned approach
- Indoor Scene Understanding in 2.5/3D for Autonomous Agents: A Survey
- R-Net: A Deep Network for Multi-oriented Vehicle Detection in Aerial Images and Videos
- Adapting Mask-RCNN for Automatic Nucleus Segmentation
- Image Augmentations for GAN Training
- Building Damage Detection in Satellite Imagery Using Convolutional Neural Networks
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training
- GhostNet: More Features from Cheap Operations
- A Comprehensive Study of Deep Video Action Recognition
- Recent Advances in Deep Learning: An Overview
- 3-D Scene Graph: A Sparse and Semantic Representation of Physical Environments for Intelligent Agents
- CLIP-Adapter: Better Vision-Language Models with Feature Adapters
- Recent Advances in Object Detection in the Age of Deep Convolutional Neural Networks
- Multivariate Confidence Calibration for Object Detection
- Convolutional Neural Networks Applied to Neutrino Events in a Liquid Argon Time Projection Chamber
- Protein Secondary Structure Prediction Using Cascaded Convolutional and Recurrent Neural Networks
- Interpreting and Improving Adversarial Robustness of Deep Neural Networks with Neuron Sensitivity
- Rethinking on Multi-Stage Networks for Human Pose Estimation
- Multi-Scale Single Image Dehazing Using Laplacian and Gaussian Pyramids
- ByteTrack: Multi-Object Tracking by Associating Every Detection Box
- Mechanical Search: Multi-Step Retrieval of a Target Object Occluded by Clutter
- Localizing Objects with Self-Supervised Transformers and no Labels
- Weakly-supervised DCNN for RGB-D Object Recognition in Real-World Applications Which Lack Large-scale Annotated Training Data
- You Only Watch Once: A Unified CNN Architecture for Real-Time Spatiotemporal Action Localization
- Bottom-up Object Detection by Grouping Extreme and Center Points
- Training Convolutional Neural Networks with Limited Training Data for Ear Recognition in the Wild
- Varifocal-Net: A Chromosome Classification Approach using Deep Convolutional Networks
- Addressing Failure Prediction by Learning Model Confidence
- Zero-Shot Detection
- Incremental Object Detection via Meta-Learning
- Siamese Attentional Keypoint Network for High Performance Visual Tracking
- Multi-task Learning with Coarse Priors for Robust Part-aware Person Re-identification
- Sparse R-CNN: End-to-End Object Detection with Learnable Proposals
- Semantics for Robotic Mapping, Perception and Interaction: A Survey
- An Empirical Study of Spatial Attention Mechanisms in Deep Networks
- Cascade R-CNN: High Quality Object Detection and Instance Segmentation
- PolarDet: A Fast, More Precise Detector for Rotated Target in Aerial Images
- Tiny-DSOD: Lightweight Object Detection for Resource-Restricted Usages
- StarNet: Targeted Computation for Object Detection in Point Clouds
- Rearrangement: A Challenge for Embodied AI
- See Better Before Looking Closer: Weakly Supervised Data Augmentation Network for Fine-Grained Visual Classification
- Pose-Aware Instance Segmentation Framework from Cone Beam CT Images for Tooth Segmentation
- Rethinking and Designing a High-performing Automatic License Plate Recognition Approach
- From Points to Parts: 3D Object Detection from Point Cloud with Part-aware and Part-aggregation Network
- Video Salient Object Detection Using Spatiotemporal Deep Features
- End-to-End Multi-View Fusion for 3D Object Detection in LiDAR Point Clouds
- A Survey on Collaborative DNN Inference for Edge Intelligence
- AU R-CNN: Encoding Expert Prior Knowledge into R-CNN for Action Unit Detection
- Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
- Instance Segmentation in the Dark
- Sewer-ML: A Multi-Label Sewer Defect Classification Dataset and Benchmark
- Learning to Fuse Things and Stuff
- Focus: Querying Large Video Datasets with Low Latency and Low Cost
- Deep Ordinal Hashing with Spatial Attention
- Abnormal Colon Polyp Image Synthesis Using Conditional Adversarial Networks for Improved Detection Performance
- Next-Active-Object prediction from Egocentric Videos
- G-Rep: Gaussian Representation for Arbitrary-Oriented Object Detection
- Finding Task-Relevant Features for Few-Shot Learning by Category Traversal
- Probabilistic and Geometric Depth: Detecting Objects in Perspective
- How Much Position Information Do Convolutional Neural Networks Encode?
- Universal Joint Feature Extraction for P300 EEG Classification using Multi-task Autoencoder
- Towards Automated Infographic Design: Deep Learning-based Auto-Extraction of Extensible Timeline
- Anomaly Detection in Traffic Scenes via Spatial-aware Motion Reconstruction
- PP-YOLOv2: A Practical Object Detector
- Self-supervised Image Enhancement Network: Training with Low Light Images Only
- Image Captioning: Transforming Objects into Words
- SaliencyMix: A Saliency Guided Data Augmentation Strategy for Better Regularization
- Accelerator-Aware Pruning for Convolutional Neural Networks
- A Deep Journey into Super-resolution: A survey
- Sparse DETR: Efficient End-to-End Object Detection with Learnable Sparsity
- Contextual Translation Embedding for Visual Relationship Detection and Scene Graph Generation
- Segmentation of cell-level anomalies in electroluminescence images of photovoltaic modules
- Adversarial Attacks Against Deep Generative Models on Data: A Survey
- Multiclass Weighted Loss for Instance Segmentation of Cluttered Cells
- Distribution Aligning Refinery of Pseudo-label for Imbalanced Semi-supervised Learning
- An application of a deep learning algorithm for automatic detection of unexpected accidents under bad CCTV monitoring conditions in tunnels
- Arbitrary-Oriented Ship Detection through Center-Head Point Extraction
- On the Origin of Deep Learning
- Perceptual Generative Adversarial Networks for Small Object Detection
- Fast Fine-grained Image Classification via Weakly Supervised Discriminative Localization
- A novel Region of Interest Extraction Layer for Instance Segmentation
- Achieving Super-Linear Speedup across Multi-FPGA for Real-Time DNN Inference
- RobustTAD: Robust Time Series Anomaly Detection via Decomposition and Convolutional Neural Networks
- AP-Loss for Accurate One-Stage Object Detection
- Scale-Aware Trident Networks for Object Detection
- CPT: Colorful Prompt Tuning for Pre-trained Vision-Language Models
- Closing the Generalization Gap of Adaptive Gradient Methods in Training Deep Neural Networks
- Object Level Visual Reasoning in Videos
- A Deep Learning System That Generates Quantitative CT Reports for Diagnosing Pulmonary Tuberculosis
- Deep Image Retrieval: Learning global representations for image search
- High precision control and deep learning-based corn stand counting algorithms for agricultural robot
- A Single-Shot Arbitrarily-Shaped Text Detector based on Context Attended Multi-Task Learning
- 3D RoI-aware U-Net for Accurate and Efficient Colorectal Tumor Segmentation
- Configurable 3D Scene Synthesis and 2D Image Rendering with Per-Pixel Ground Truth using Stochastic Grammars
- Stacked Deconvolutional Network for Semantic Segmentation
- What makes instance discrimination good for transfer learning?
- Recurrent Video Deblurring with Blur-Invariant Motion Estimation and Pixel Volumes
- Spatiotemporal Contrastive Video Representation Learning
- Deep Divergence-Based Approach to Clustering
- Exploring Randomly Wired Neural Networks for Image Recognition
- FIgLib & SmokeyNet: Dataset and Deep Learning Model for Real-Time Wildland Fire Smoke Detection
- Towards Robust LiDAR-based Perception in Autonomous Driving: General Black-box Adversarial Sensor Attack and Countermeasures
- Advances and Applications of Computer Vision Techniques in Vehicle Trajectory Generation and Surrogate Traffic Safety Indicators
- Deep Learning Inference in Facebook Data Centers: Characterization, Performance Optimizations and Hardware Implications
- Learning Rich Features for Image Manipulation Detection
- DirectPose: Direct End-to-End Multi-Person Pose Estimation
- Tencent ML-Images: A Large-Scale Multi-Label Image Database for Visual Representation Learning
- Deep Learning Techniques for In-Crop Weed Identification: A Review
- SAR Image Despeckling by Deep Neural Networks: from a pre-trained model to an end-to-end training strategy
- Online PCB Defect Detector On A New PCB Defect Dataset
- Occlusion Handling in Generic Object Detection: A Review
- Local Aggregation for Unsupervised Learning of Visual Embeddings
- Improving Transferability of Adversarial Examples with Input Diversity
- Soft Rasterizer: Differentiable Rendering for Unsupervised Single-View Mesh Reconstruction
- Supervised Compression for Resource-Constrained Edge Computing Systems
- A Multi-Level Approach to Waste Object Segmentation
- Complex-YOLO: Real-time 3D Object Detection on Point Clouds
- InsPLAD: A Dataset and Benchmark for Power Line Asset Inspection in UAV Images
- HigherHRNet: Scale-Aware Representation Learning for Bottom-Up Human Pose Estimation
- IENet: Interacting Embranchment One Stage Anchor Free Detector for Orientation Aerial Object Detection
- Multi-Class Multi-Object Tracking using Changing Point Detection
- Weakly-Supervised Video Object Grounding from Text by Loss Weighting and Object Interaction
- Enabling Pedestrian Safety using Computer Vision Techniques: A Case Study of the 2018 Uber Inc. Self-driving Car Crash
- Efficient and Accurate Arbitrary-Shaped Text Detection with Pixel Aggregation Network
- AutoDNNchip: An Automated DNN Chip Predictor and Builder for Both FPGAs and ASICs
- Language-guided Navigation via Cross-Modal Grounding and Alternate Adversarial Learning
- Handgun detection using combined human pose and weapon appearance
- Depth Adaptive Deep Neural Network for Semantic Segmentation
- MeliusNet: Can Binary Neural Networks Achieve MobileNet-level Accuracy?
- Traffic Light Recognition Using Deep Learning and Prior Maps for Autonomous Cars
- Learning Multi-Attention Context Graph for Group-Based Re-Identification
- Transferable Interactiveness Knowledge for Human-Object Interaction Detection
- Vehicle Tracking Using Surveillance with Multimodal Data Fusion
- Visual and Semantic Knowledge Transfer for Large Scale Semi-supervised Object Detection
- Towards seamless multi-view scene analysis from satellite to street-level
- A Multi-cut Formulation for Joint Segmentation and Tracking of Multiple Objects
- RepPoints V2: Verification Meets Regression for Object Detection
- Few-Shot Learning with Geometric Constraints
- Driver Action Prediction Using Deep (Bidirectional) Recurrent Neural Network
- Transfer Learning-based Road Damage Detection for Multiple Countries
- Artificial Neural Network Based Breast Cancer Screening: A Comprehensive Review
- Pseudo-LiDAR from Visual Depth Estimation: Bridging the Gap in 3D Object Detection for Autonomous Driving
- BERTgrid: Contextualized Embedding for 2D Document Representation and Understanding
- TinaFace: Strong but Simple Baseline for Face Detection
- Attention-guided Context Feature Pyramid Network for Object Detection
- Center-based 3D Object Detection and Tracking
- EfficientPose: An efficient, accurate and scalable end-to-end 6D multi object pose estimation approach
- Relation Networks for Object Detection
- Image Captioning for Effective Use of Language Models in Knowledge-Based Visual Question Answering
- Self-EMD: Self-Supervised Object Detection without ImageNet
- Bilinear CNNs for Fine-grained Visual Recognition
- Optimizing Rank-based Metrics with Blackbox Differentiation
- A Study of Face Obfuscation in ImageNet
- RiFCN: Recurrent Network in Fully Convolutional Network for Semantic Segmentation of High Resolution Remote Sensing Images
- One-Shot Instance Segmentation
- Bayesian Loss for Crowd Count Estimation with Point Supervision
- Monocular 3D Object Detection and Box Fitting Trained End-to-End Using Intersection-over-Union Loss
- End-to-End Video Instance Segmentation with Transformers
- Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign Dropout
- HOI Analysis: Integrating and Decomposing Human-Object Interaction
- Face Detection Using Improved Faster RCNN
- Accelerating Flash Calculation through Deep Learning Methods
- Rethinking the Hyperparameters for Fine-tuning
- Learning Modulated Loss for Rotated Object Detection
- Neural Compression and Filtering for Edge-assisted Real-time Object Detection in Challenged Networks
- MDSSD: Multi-scale Deconvolutional Single Shot Detector for Small Objects
- Key Points Estimation and Point Instance Segmentation Approach for Lane Detection
- Learning to Refine Object Segments
- SOLVER: Scene-Object Interrelated Visual Emotion Reasoning Network
- ASFormer: Transformer for Action Segmentation
- Deep Set Prediction Networks
- Neural Attention for Image Captioning: Review of Outstanding Methods
- Panoptic Feature Pyramid Networks
- Domain Adaptation for Object Detection via Style Consistency
- MAMBA: Multi-level Aggregation via Memory Bank for Video Object Detection
- Understanding Straight-Through Estimator in Training Activation Quantized Neural Nets
- ConvTransformer: A Convolutional Transformer Network for Video Frame Synthesis
- Empowering Things with Intelligence: A Survey of the Progress, Challenges, and Opportunities in Artificial Intelligence of Things
- Multi-modal Deep Analysis for Multimedia
- Damage detection using in-domain and cross-domain transfer learning
- Security and Privacy Issues in Deep Learning
- Cascaded Boundary Regression for Temporal Action Detection
- Region Proposal by Guided Anchoring
- Learning Human-Object Interactions by Graph Parsing Neural Networks
- One-Two-One Networks for Compression Artifacts Reduction in Remote Sensing
- Finding any Waldo: zero-shot invariant and efficient visual search
- Radar Voxel Fusion for 3D Object Detection
- Relation-Aware Graph Attention Network for Visual Question Answering
- Data Augmentation for Object Detection via Progressive and Selective Instance-Switching
- InterBERT: Vision-and-Language Interaction for Multi-modal Pretraining
- Apple Counting using Convolutional Neural Networks
- SOLQ: Segmenting Objects by Learning Queries
- Multiscale Vision Transformers
- Resolving Class Imbalance in Object Detection with Weighted Cross Entropy Losses
- Overcoming Small Minirhizotron Datasets Using Transfer Learning
- Privacy Protection in Street-View Panoramas using Depth and Multi-View Imagery
- Forecasting Action through Contact Representations from First Person Video
- Driving Scene Perception Network: Real-time Joint Detection, Depth Estimation and Semantic Segmentation
- i-Mix: A Domain-Agnostic Strategy for Contrastive Representation Learning
- Lightweight Modules for Efficient Deep Learning based Image Restoration
- Normalized Cut Loss for Weakly-supervised CNN Segmentation
- Pyramid Mask Text Detector
- BlendMask: Top-Down Meets Bottom-Up for Instance Segmentation
- Uncertainty for Identifying Open-Set Errors in Visual Object Detection
- Background Learnable Cascade for Zero-Shot Object Detection
- Deep Learning for Plasma Tomography and Disruption Prediction from Bolometer Data
- PixelLink: Detecting Scene Text via Instance Segmentation
- DEFT: Detection Embeddings for Tracking
- Soft Sampling for Robust Object Detection
- Tensor Programs II: Neural Tangent Kernel for Any Architecture
- SS-CAM: Smoothed Score-CAM for Sharper Visual Feature Localization
- Small-scale Pedestrian Detection Based on Somatic Topology Localization and Temporal Feature Aggregation
- Real Time Bangladeshi Sign Language Detection using Faster R-CNN
- STN-OCR: A single Neural Network for Text Detection and Text Recognition
- A Comparison of CNN-based Face and Head Detectors for Real-Time Video Surveillance Applications
- Object DGCNN: 3D Object Detection using Dynamic Graphs
- Progressive Coordinate Transforms for Monocular 3D Object Detection
- Toward unsupervised, multi-object discovery in large-scale image collections
- Cascade RetinaNet: Maintaining Consistency for Single-Stage Object Detection
- Adversarial Examples in Modern Machine Learning: A Review
- On the Binding Problem in Artificial Neural Networks
- Simultaneously Localize, Segment and Rank the Camouflaged Objects
- 3D Object Proposals using Stereo Imagery for Accurate Object Class Detection
- Joint Monocular 3D Vehicle Detection and Tracking
- Trends in Integration of Vision and Language Research: A Survey of Tasks, Datasets, and Methods
- How convolutional neural network see the world - A survey of convolutional neural network visualization methods
- Efficient Neural Architecture Transformation Searchin Channel-Level for Object Detection
- Reference-based Defect Detection Network
- Self-supervising Action Recognition by Statistical Moment and Subspace Descriptors
- MEAL V2: Boosting Vanilla ResNet-50 to 80%+ Top-1 Accuracy on ImageNet without Tricks
- Less is More: ClipBERT for Video-and-Language Learning via Sparse Sampling
- VideoMix: Rethinking Data Augmentation for Video Classification
- M6: A Chinese Multimodal Pretrainer
- Syn2Real: A New Benchmark forSynthetic-to-Real Visual Domain Adaptation
- CFUN: Combining Faster R-CNN and U-net Network for Efficient Whole Heart Segmentation
- Spiking-YOLO: Spiking Neural Network for Energy-Efficient Object Detection
- Stacked Cross Attention for Image-Text Matching
- Self-similarity Grouping: A Simple Unsupervised Cross Domain Adaptation Approach for Person Re-identification
- Spatio-Temporal Point Process for Multiple Object Tracking
- Model Rubik's Cube: Twisting Resolution, Depth and Width for TinyNets
- Early Detection of Retinopathy of Prematurity stage using Deep Learning approach
- CenterNet3D: An Anchor Free Object Detector for Point Cloud
- ViDT: An Efficient and Effective Fully Transformer-based Object Detector
- Traffic-Aware Multi-Camera Tracking of Vehicles Based on ReID and Camera Link Model
- LatentGNN: Learning Efficient Non-local Relations for Visual Recognition
- Robust, Occlusion-aware Pose Estimation for Objects Grasped by Adaptive Hands
- Knowledge-Embedded Routing Network for Scene Graph Generation
- Visual Object Tracking in First Person Vision
- PDNet: Toward Better One-Stage Object Detection With Prediction Decoupling
- Back to Simplicity: How to Train Accurate BNNs from Scratch?
- Distilling Object Detectors with Task Adaptive Regularization
- Detecting Multi-Oriented Text with Corner-based Region Proposals
- TNCR: Table Net Detection and Classification Dataset
- Deep Weakly-Supervised Learning Methods for Classification and Localization in Histology Images: A Survey
- On the detection-to-track association for online multi-object tracking
- CARAFE: Content-Aware ReAssembly of FEatures
- Prime Sample Attention in Object Detection
- Defending Against Universal Attacks Through Selective Feature Regeneration
- Decoupled Classification Refinement: Hard False Positive Suppression for Object Detection
- DeepIris: Iris Recognition Using A Deep Learning Approach
- Towards Unconstrained Palmprint Recognition on Consumer Devices: a Literature Review
- The QXS-SAROPT Dataset for Deep Learning in SAR-Optical Data Fusion
- Combination of Multiple Global Descriptors for Image Retrieval
- Visual Affordance and Function Understanding: A Survey
- Multi-Target Deep Learning for Algal Detection and Classification
- Object Detection for Comics using Manga109 Annotations
- Comprehensive Instructional Video Analysis: The COIN Dataset and Performance Evaluation
- LFFD: A Light and Fast Face Detector for Edge Devices
- AugFPN: Improving Multi-scale Feature Learning for Object Detection
- Tool Detection and Operative Skill Assessment in Surgical Videos Using Region-Based Convolutional Neural Networks
- CrowdPose: Efficient Crowded Scenes Pose Estimation and A New Benchmark
- Multi-Channel CNN-based Object Detection for Enhanced Situation Awareness
- S3D: Single Shot multi-Span Detector via Fully 3D Convolutional Networks
- Position, Padding and Predictions: A Deeper Look at Position Information in CNNs
- DeLTA: GPU Performance Model for Deep Learning Applications with In-depth Memory System Traffic Analysis
- Panoptic-DeepLab: A Simple, Strong, and Fast Baseline for Bottom-Up Panoptic Segmentation
- Hearing What You Cannot See: Acoustic Vehicle Detection Around Corners
- Multi-adversarial Faster-RCNN for Unrestricted Object Detection
- Human in Events: A Large-Scale Benchmark for Human-centric Video Analysis in Complex Events
- Geometric Pose Affordance: 3D Human Pose with Scene Constraints
- Bounding Box Regression with Uncertainty for Accurate Object Detection
- Distilling Knowledge by Mimicking Features
- CLOCs: Camera-LiDAR Object Candidates Fusion for 3D Object Detection
- Detection of 3D Bounding Boxes of Vehicles Using Perspective Transformation for Accurate Speed Measurement
- ExpandNets: Linear Over-parameterization to Train Compact Convolutional Networks
- SiamCAR: Siamese Fully Convolutional Classification and Regression for Visual Tracking
- MobileDets: Searching for Object Detection Architectures for Mobile Accelerators
- QueryProp: Object Query Propagation for High-Performance Video Object Detection
- Towards End-to-End Car License Plates Detection and Recognition with Deep Neural Networks
- A Recurrent Vision-and-Language BERT for Navigation
- Voxel-FPN: multi-scale voxel feature aggregation in 3D object detection from point clouds
- PubLayNet: largest dataset ever for document layout analysis
- DADA: Differentiable Automatic Data Augmentation
- Dense Contrastive Learning for Self-Supervised Visual Pre-Training
- A Vision-based Social Distancing and Critical Density Detection System for COVID-19
- Improved training of binary networks for human pose estimation and image recognition
- FcaNet: Frequency Channel Attention Networks
- Online Structured Sparsity-based Moving Object Detection from Satellite Videos
- Physical Adversarial Attack on Vehicle Detector in the Carla Simulator
- AdaScale: Towards Real-time Video Object Detection Using Adaptive Scaling
- Unbiased Scene Graph Generation from Biased Training
- HAKE: Human Activity Knowledge Engine
- Single-shot 3D multi-person pose estimation in complex images
- ResizeMix: Mixing Data with Preserved Object Information and True Labels
- AMMUS : A Survey of Transformer-based Pretrained Models in Natural Language Processing
- Sequence Level Semantics Aggregation for Video Object Detection
- Anchor Cascade for Efficient Face Detection
- FingerNet: Pushing The Limits of Fingerprint Recognition Using Convolutional Neural Network
- Tracklet Association Tracker: An End-to-End Learning-based Association Approach for Multi-Object Tracking
- Fully Convolutional Networks for Chip-wise Defect Detection Employing Photoluminescence Images
- Class-Difficulty Based Methods for Long-Tailed Visual Recognition
- FCOS: A simple and strong anchor-free object detector
- 3D Multi-Object Tracking: A Baseline and New Evaluation Metrics
- A Robust Illumination-Invariant Camera System for Agricultural Applications
- Graph-based Knowledge Distillation by Multi-head Attention Network
- Text-Based Person Search with Limited Data
- DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting
- SOLO: Segmenting Objects by Locations
- CDSA: Cross-Dimensional Self-Attention for Multivariate, Geo-tagged Time Series Imputation
- Unifying Multimodal Transformer for Bi-directional Image and Text Generation
- Few-Shot Object Detection with Attention-RPN and Multi-Relation Detector
- Context Model for Pedestrian Intention Prediction using Factored Latent-Dynamic Conditional Random Fields
- Shape Robust Text Detection with Progressive Scale Expansion Network
- ABCNet: Real-time Scene Text Spotting with Adaptive Bezier-Curve Network
- PyramidBox++: High Performance Detector for Finding Tiny Face
- Compositional Memory for Visual Question Answering
- WoodFisher: Efficient Second-Order Approximation for Neural Network Compression
- Image-based table recognition: data, model, and evaluation
- Conditional Gaussian Distribution Learning for Open Set Recognition
- ChaLearn Looking at People: IsoGD and ConGD Large-scale RGB-D Gesture Recognition
- Saccader: Improving Accuracy of Hard Attention Models for Vision
- Fast object detection in compressed JPEG Images
- RelationNet++: Bridging Visual Representations for Object Detection via Transformer Decoder
- Learning non-maximum suppression
- Monocular 3D Object Detection with Sequential Feature Association and Depth Hint Augmentation
- Learning to Transfer Examples for Partial Domain Adaptation
- TensorMask: A Foundation for Dense Object Segmentation
- Cut-Thumbnail: A Novel Data Augmentation for Convolutional Neural Network
- Mask-Guided Attention Network for Occluded Pedestrian Detection
- UG Track 2: A Collective Benchmark Effort for Evaluating and Advancing Image Understanding in Poor Visibility Environments
- Dual-stream Network for Visual Recognition
- Distilling Object Detectors with Feature Richness
- EXTD: Extremely Tiny Face Detector via Iterative Filter Reuse
- Mining the Benefits of Two-stage and One-stage HOI Detection
- FastPose: Towards Real-time Pose Estimation and Tracking via Scale-normalized Multi-task Networks
- Heterogeneous Memory Enhanced Multimodal Attention Model for Video Question Answering
- Data Augmentation Revisited: Rethinking the Distribution Gap between Clean and Augmented Data
- SSA-CNN: Semantic Self-Attention CNN for Pedestrian Detection
- Deformable Kernels: Adapting Effective Receptive Fields for Object Deformation
- Tracklets Predicting Based Adaptive Graph Tracking
- A Multimodal Memes Classification: A Survey and Open Research Issues
- CPS++: Improving Class-level 6D Pose and Shape Estimation From Monocular Images With Self-Supervised Learning
- CFENet: An Accurate and Efficient Single-Shot Object Detector for Autonomous Driving
- Objects detection for remote sensing images based on polar coordinates
- Learning Human-Object Interaction Detection using Interaction Points
- A Learning Framework for n-bit Quantized Neural Networks toward FPGAs
- P4Contrast: Contrastive Learning with Pairs of Point-Pixel Pairs for RGB-D Scene Understanding
- An original framework for Wheat Head Detection using Deep, Semi-supervised and Ensemble Learning within Global Wheat Head Detection (GWHD) Dataset
- Attention Branch Network: Learning of Attention Mechanism for Visual Explanation
- Early Action Prediction with Generative Adversarial Networks
- Multimodal Unified Attention Networks for Vision-and-Language Interactions
- 6-DoF Object Pose from Semantic Keypoints
- EmbedMask: Embedding Coupling for One-stage Instance Segmentation
- Autonomous Driving with Deep Learning: A Survey of State-of-Art Technologies
- f-CNN: A Toolflow for Mapping Multi-CNN Applications on FPGAs
- PatchUp: A Feature-Space Block-Level Regularization Technique for Convolutional Neural Networks
- Training Binary Neural Networks through Learning with Noisy Supervision
- An Empirical Study on Leveraging Scene Graphs for Visual Question Answering
- Multi-Task Incremental Learning for Object Detection
- Data Extraction from Charts via Single Deep Neural Network
- Fast Visual Object Tracking with Rotated Bounding Boxes
- SFD: Single Shot Scale-invariant Face Detector
- Exploring Self-attention for Image Recognition
- Hybrid Channel Based Pedestrian Detection
- AP-MTL: Attention Pruned Multi-task Learning Model for Real-time Instrument Detection and Segmentation in Robot-assisted Surgery
- End-to-End Incremental Learning
- Object Detection Networks on Convolutional Feature Maps
- Learning Independent Instance Maps for Crowd Localization
- Negative Margin Matters: Understanding Margin in Few-shot Classification
- Long-Term Feature Banks for Detailed Video Understanding
- Extending Maps with Semantic and Contextual Object Information for Robot Navigation: a Learning-Based Framework using Visual and Depth Cues
- Stroke-Based Scene Text Erasing Using Synthetic Data for Training
- Virtual to Real adaptation of Pedestrian Detectors
- VarifocalNet: An IoU-aware Dense Object Detector
- Locate, Size and Count: Accurately Resolving People in Dense Crowds via Detection
- Learning Spatiotemporal Features via Video and Text Pair Discrimination
- D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual Odometry
- Robustness of Object Recognition under Extreme Occlusion in Humans and Computational Models
- Where are the Masks: Instance Segmentation with Image-level Supervision
- Variational Context: Exploiting Visual and Textual Context for Grounding Referring Expressions
- SODA10M: A Large-Scale 2D Self/Semi-Supervised Object Detection Dataset for Autonomous Driving
- Deep Neural Network with l2-norm Unit for Brain Lesions Detection
- Classification based Grasp Detection using Spatial Transformer Network
- TextSnake: A Flexible Representation for Detecting Text of Arbitrary Shapes
- Diversify and Match: A Domain Adaptive Representation Learning Paradigm for Object Detection
- PointGroup: Dual-Set Point Grouping for 3D Instance Segmentation
- An embedded system for the automated generation of labeled plant images to enable machine learning applications in agriculture
- Learning in an Uncertain World: Representing Ambiguity Through Multiple Hypotheses
- Deep Texture-Aware Features for Camouflaged Object Detection
- Deep Co-supervision and Attention Fusion Strategy for Automatic COVID-19 Lung Infection Segmentation on CT Images
- Improved Selective Refinement Network for Face Detection
- Towards Optimal Structured CNN Pruning via Generative Adversarial Learning
- Radar-Camera Sensor Fusion for Joint Object Detection and Distance Estimation in Autonomous Vehicles
- TableSense: Spreadsheet Table Detection with Convolutional Neural Networks
- Concurrent Activity Recognition with Multimodal CNN-LSTM Structure
- Chained-Tracker: Chaining Paired Attentive Regression Results for End-to-End Joint Multiple-Object Detection and Tracking
- Gradient Centralization: A New Optimization Technique for Deep Neural Networks
- MaX-DeepLab: End-to-End Panoptic Segmentation with Mask Transformers
- Towards Real-Time Multi-Object Tracking
- DeepFL-IQA: Weak Supervision for Deep IQA Feature Learning
- Sign Language Recognition via Skeleton-Aware Multi-Model Ensemble
- Benchmarking Single Image Dehazing and Beyond
- AIBench: An Industry Standard Internet Service AI Benchmark Suite
- Generation of microbial colonies dataset with deep learning style transfer
- Feature Fusion for Online Mutual Knowledge Distillation
- The Lottery Tickets Hypothesis for Supervised and Self-supervised Pre-training in Computer Vision Models
- 3D IoU-Net: IoU Guided 3D Object Detector for Point Clouds
- It Takes Two to Tango: Towards Theory of AI's Mind
- Exploring Object Relation in Mean Teacher for Cross-Domain Detection
- A Shape Transformation-based Dataset Augmentation Framework for Pedestrian Detection
- Are we pretraining it right? Digging deeper into visio-linguistic pretraining
- C-MIL: Continuation Multiple Instance Learning for Weakly Supervised Object Detection
- SSAP: Single-Shot Instance Segmentation With Affinity Pyramid
- Occlusion-aware R-CNN: Detecting Pedestrians in a Crowd
- Low-light Image Enhancement Algorithm Based on Retinex and Generative Adversarial Network
- Learning Deep ResNet Blocks Sequentially using Boosting Theory
- VD-BERT: A Unified Vision and Dialog Transformer with BERT
- A Deep One-Shot Network for Query-based Logo Retrieval
- SIFT Meets CNN: A Decade Survey of Instance Retrieval
- Looking Fast and Slow: Memory-Guided Mobile Video Object Detection
- Learning to Assemble Neural Module Tree Networks for Visual Grounding
- Video Action Understanding
- Multimodal Transformer with Multi-View Visual Representation for Image Captioning
- DeepBox: Learning Objectness with Convolutional Networks
- Scene Graph Generation with External Knowledge and Image Reconstruction
- Memory Enhanced Global-Local Aggregation for Video Object Detection
- Multiple Object Tracking by Flowing and Fusing
- Relation Distillation Networks for Video Object Detection
- TANet: Robust 3D Object Detection from Point Clouds with Triple Attention
- Dynamic R-CNN: Towards High Quality Object Detection via Dynamic Training
- Graph Density-Aware Losses for Novel Compositions in Scene Graph Generation
- Rethinking Class-Balanced Methods for Long-Tailed Visual Recognition from a Domain Adaptation Perspective
- CentripetalNet: Pursuing High-quality Keypoint Pairs for Object Detection
- Improving Auto-Augment via Augmentation-Wise Weight Sharing
- An Annotation Saved is an Annotation Earned: Using Fully Synthetic Training for Object Instance Detection
- SRG: Snippet Relatedness-based Temporal Action Proposal Generator
- MiniVLM: A Smaller and Faster Vision-Language Model
- CycleSegNet: Object Co-segmentation with Cycle Refinement and Region Correspondence
- FSCE: Few-Shot Object Detection via Contrastive Proposal Encoding
- PV-RCNN++: Point-Voxel Feature Set Abstraction With Local Vector Representation for 3D Object Detection
- PPDM: Parallel Point Detection and Matching for Real-time Human-Object Interaction Detection
- FedVision: An Online Visual Object Detection Platform Powered by Federated Learning
- Efficient Bird Eye View Proposals for 3D Siamese Tracking
- Pillar-based Object Detection for Autonomous Driving
- Tracking Holistic Object Representations
- Age and Gender Prediction From Face Images Using Attentional Convolutional Network
- Scaling Wide Residual Networks for Panoptic Segmentation
- Simultaneous multi-view instance detection with learned geometric soft-constraints
- G-TAD: Sub-Graph Localization for Temporal Action Detection
- A Survey of Modern Object Detection Literature using Deep Learning
- Weight-Sharing Neural Architecture Search: A Battle to Shrink the Optimization Gap
- Multi-step Reasoning via Recurrent Dual Attention for Visual Dialog
- MLCVNet: Multi-Level Context VoteNet for 3D Object Detection
- Improving Image Captioning with Better Use of Captions
- UFO: A UniFied TransfOrmer for Vision-Language Representation Learning
- One-shot Face Reenactment
- Exploring Categorical Regularization for Domain Adaptive Object Detection
- Universal Adversarial Perturbations: A Survey
- WIDER Face and Pedestrian Challenge 2018: Methods and Results
- BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video Analytics
- Large-Scale Generative Data-Free Distillation
- Visual Relationship Detection with Visual-Linguistic Knowledge from Multimodal Representations
- Dynamic Adversarial Patch for Evading Object Detection Models
- Decoupled Adaptation for Cross-Domain Object Detection
- Towards End-to-end Text Spotting with Convolutional Recurrent Neural Networks
- The H3D Dataset for Full-Surround 3D Multi-Object Detection and Tracking in Crowded Urban Scenes
- FCOS3D: Fully Convolutional One-Stage Monocular 3D Object Detection
- Spatial-Temporal Relation Networks for Multi-Object Tracking
- On Feature Normalization and Data Augmentation
- Chainer: A Deep Learning Framework for Accelerating the Research Cycle
- InstaBoost: Boosting Instance Segmentation via Probability Map Guided Copy-Pasting
- FAMNet: Joint Learning of Feature, Affinity and Multi-dimensional Assignment for Online Multiple Object Tracking
- Language-Conditioned Graph Networks for Relational Reasoning
- MobilePose: Real-Time Pose Estimation for Unseen Objects with Weak Shape Supervision
- Improving Object Detection from Scratch via Gated Feature Reuse
- Towards a Robust Deep Neural Network in Texts: A Survey
- Simple Training Strategies and Model Scaling for Object Detection
- Light-Weight RetinaNet for Object Detection
- Cell Detection in Microscopy Images with Deep Convolutional Neural Network and Compressed Sensing
- Weakly Supervised Dense Video Captioning
- Reduced Focal Loss: 1st Place Solution to xView object detection in Satellite Imagery
- Forward and Backward Information Retention for Accurate Binary Neural Networks
- A benchmark dataset for deep learning-based airplane detection: HRPlanes
- Conditional Convolutions for Instance Segmentation
- AffordanceNet: An End-to-End Deep Learning Approach for Object Affordance Detection
- Detecting and Identifying Optical Signal Attacks on Autonomous Driving Systems
- Image Amodal Completion: A Survey
- Learning Multi-level Deep Representations for Image Emotion Classification
- Deep Co-Training for Semi-Supervised Image Recognition
- A Survey on Deep Learning Toolkits and Libraries for Intelligent User Interfaces
- Unpaired Image Captioning via Scene Graph Alignments
- A Survey of Deep Learning Approaches for OCR and Document Understanding
- Normalized and Geometry-Aware Self-Attention Network for Image Captioning
- VID-WIN: Fast Video Event Matching with Query-Aware Windowing at the Edge for the Internet of Multimedia Things
- Learning Accurate Low-Bit Deep Neural Networks with Stochastic Quantization
- Text Detection and Recognition in the Wild: A Review
- Real-Time and Accurate Object Detection in Compressed Video by Long Short-term Feature Aggregation
- Distilling Object Detectors via Decoupled Features
- YOLObile: Real-Time Object Detection on Mobile Devices via Compression-Compilation Co-Design
- HARRISON: A Benchmark on HAshtag Recommendation for Real-world Images in Social Networks
- Fourier Contour Embedding for Arbitrary-Shaped Text Detection
- 3D Point Cloud Descriptors in Hand-crafted and Deep Learning Age: State-of-the-Art
- DBF: Dynamic Belief Fusion for Combining Multiple Object Detectors
- Rethinking "Batch" in BatchNorm
- Exploring Data Augmentation for Multi-Modality 3D Object Detection
- Preferences Prediction using a Gallery of Mobile Device based on Scene Recognition and Object Detection
- End-to-end Flow Correlation Tracking with Spatial-temporal Attention
- Deep Neural Networks for Marine Debris Detection in Sonar Images
- Single-Stage Multi-Person Pose Machines
- Progressive Sparse Local Attention for Video object detection
- e-SNLI-VE: Corrected Visual-Textual Entailment with Natural Language Explanations
- Understanding the Behaviour of Contrastive Loss
- Generalized Focal Loss V2: Learning Reliable Localization Quality Estimation for Dense Object Detection
- Computing Systems for Autonomous Driving: State-of-the-Art and Challenges
- DeepFashion2: A Versatile Benchmark for Detection, Pose Estimation, Segmentation and Re-Identification of Clothing Images
- Direction Concentration Learning: Enhancing Congruency in Machine Learning
- Restoring Negative Information in Few-Shot Object Detection
- Weakly Aligned Cross-Modal Learning for Multispectral Pedestrian Detection
- A Performance Comparison of Loss Functions for Deep Face Recognition
- Objectness Scoring and Detection Proposals in Forward-Looking Sonar Images with Convolutional Neural Networks
- Corner Proposal Network for Anchor-free, Two-stage Object Detection
- LAEO-Net++: revisiting people Looking At Each Other in videos
- Recent Advances in Deep Learning for Object Detection
- Autonomous and cooperative design of the monitor positions for a team of UAVs to maximize the quantity and quality of detected objects
- Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning
- Involution: Inverting the Inherence of Convolution for Visual Recognition
- Deep Learning for LiDAR Point Clouds in Autonomous Driving: A Review
- Attention Based Glaucoma Detection: A Large-scale Database and CNN Model
- Zoom Out-and-In Network with Recursive Training for Object Proposal
- MITOS-RCNN: A Novel Approach to Mitotic Figure Detection in Breast Cancer Histopathology Images using Region Based Convolutional Neural Networks
- How To Train Your Deep Multi-Object Tracker
- CSL-YOLO: A New Lightweight Object Detection System for Edge Computing
- Deep Contextual Attention for Human-Object Interaction Detection
- An explainable deep vision system for animal classification and detection in trail-camera images with automatic post-deployment retraining
- Dynamic Anchor Learning for Arbitrary-Oriented Object Detection
- Incremental Deep Learning for Robust Object Detection in Unknown Cluttered Environments
- PAN++: Towards Efficient and Accurate End-to-End Spotting of Arbitrarily-Shaped Text
- Distribution Alignment: A Unified Framework for Long-tail Visual Recognition
- Revisiting Locally Supervised Learning: an Alternative to End-to-end Training
- Attention-guided Unified Network for Panoptic Segmentation
- The More You Know: Using Knowledge Graphs for Image Classification
- GradAug: A New Regularization Method for Deep Neural Networks
- Domain Adaptation without Source Data
- Multi-Object Tracking with Siamese Track-RCNN
- SA-Net: Deep Neural Network for Robot Trajectory Recognition from RGB-D Streams
- From Human Mesenchymal Stromal Cells to Osteosarcoma Cells Classification by Deep Learning
- Detection in Crowded Scenes: One Proposal, Multiple Predictions
- Spatially Adaptive Computation Time for Residual Networks
- Spatio-Temporal Attention Models for Grounded Video Captioning
- Meta-DETR: Image-Level Few-Shot Object Detection with Inter-Class Correlation Exploitation
- GoDP: Globally optimized dual pathway system for facial landmark localization in-the-wild
- Adaptive NMS: Refining Pedestrian Detection in a Crowd
- Dual-Level Collaborative Transformer for Image Captioning
- Learning Lightweight Pedestrian Detector with Hierarchical Knowledge Distillation
- Person Search via A Mask-Guided Two-Stream CNN Model
- Crafting GBD-Net for Object Detection
- PCONV: The Missing but Desirable Sparsity in DNN Weight Pruning for Real-time Execution on Mobile Devices
- Towards Balanced Learning for Instance Recognition
- Propagate Yourself: Exploring Pixel-Level Consistency for Unsupervised Visual Representation Learning
- Equalization Loss for Long-Tailed Object Recognition
- X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal Transformers
- Extended Feature Pyramid Network for Small Object Detection
- X-volution: On the unification of convolution and self-attention
- Mixing Real and Synthetic Data to Enhance Neural Network Training -- A Review of Current Approaches
- DP-Image: Differential Privacy for Image Data in Feature Space
- AutoSweep: Recovering 3D Editable Objectsfrom a Single Photograph
- Harmonizing Transferability and Discriminability for Adapting Object Detectors
- Egocentric Image Captioning for Privacy-Preserved Passive Dietary Intake Monitoring
- Locally Free Weight Sharing for Network Width Search
- Disentangling Label Distribution for Long-tailed Visual Recognition
- ConFoc: Content-Focus Protection Against Trojan Attacks on Neural Networks
- Regularizing Deep Networks with Semantic Data Augmentation
- Towards Unconstrained End-to-End Text Spotting
- Multiple instance learning on deep features for weakly supervised object detection with extreme domain shifts
- NAS-FCOS: Fast Neural Architecture Search for Object Detection
- Image Segmentation via Probabilistic Graph Matching
- A Light Dual-Task Neural Network for Haze Removal
- Video action detection by learning graph-based spatio-temporal interactions
- Fixing the train-test resolution discrepancy
- Deep Reinforcement Learning in Computer Vision: A Comprehensive Survey
- Learning the Best Pooling Strategy for Visual Semantic Embedding
- Simple Unsupervised Multi-Object Tracking
- Object as Hotspots: An Anchor-Free 3D Object Detection Approach via Firing of Hotspots
- Improving RetinaNet for CT Lesion Detection with Dense Masks from Weak RECIST Labels
- Hiding Faces in Plain Sight: Disrupting AI Face Synthesis with Adversarial Perturbations
- Visual Semantic Reasoning for Image-Text Matching
- Context-Driven Detection of Invertebrate Species in Deep-Sea Video
- Training Deep Learning Based Denoisers without Ground Truth Data
- Automated System for Ship Detection from Medium Resolution Satellite Optical Imagery
- Say As You Wish: Fine-grained Control of Image Caption Generation with Abstract Scene Graphs
- EdgeStereo: A Context Integrated Residual Pyramid Network for Stereo Matching
- Deep Learning for Scene Classification: A Survey
- Towards Compact and Robust Deep Neural Networks
- An End-to-End Network for Panoptic Segmentation
- Transformer-Based Source-Free Domain Adaptation
- Multi-Person Pose Estimation with Local Joint-to-Person Associations
- Pointly-Supervised Action Localization
- Active and Incremental Learning with Weak Supervision
- SIXray : A Large-scale Security Inspection X-ray Benchmark for Prohibited Item Discovery in Overlapping Images
- Monocular 3D Object Detection Leveraging Accurate Proposals and Shape Reconstruction
- A Real-Time Cross-modality Correlation Filtering Method for Referring Expression Comprehension
- Language-Conditioned Imitation Learning for Robot Manipulation Tasks
- Towards Practical Multi-Object Manipulation using Relational Reinforcement Learning
- A2J: Anchor-to-Joint Regression Network for 3D Articulated Pose Estimation from a Single Depth Image
- QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object Detection
- Multimodal Integration of Human-Like Attention in Visual Question Answering
- ShapeMask: Learning to Segment Novel Objects by Refining Shape Priors
- Pneumothorax Segmentation: Deep Learning Image Segmentation to predict Pneumothorax
- Embodied Visual Recognition
- Attentional Network for Visual Object Detection
- Multi-Scale Positive Sample Refinement for Few-Shot Object Detection
- MOPT: Multi-Object Panoptic Tracking
- Enhancing Human Pose Estimation in Ancient Vase Paintings via Perceptually-grounded Style Transfer Learning
- 3D-SIS: 3D Semantic Instance Segmentation of RGB-D Scans
- ABCNet v2: Adaptive Bezier-Curve Network for Real-time End-to-end Text Spotting
- Temporally Grounding Language Queries in Videos by Contextual Boundary-aware Prediction
- Dog Identification using Soft Biometrics and Neural Networks
- Focal and Global Knowledge Distillation for Detectors
- An Accurate and Real-time Self-blast Glass Insulator Location Method Based On Faster R-CNN and U-net with Aerial Images
- Adapting Object Detectors with Conditional Domain Normalization
- Practical Detection of Trojan Neural Networks: Data-Limited and Data-Free Cases
- Learning Densities in Feature Space for Reliable Segmentation of Indoor Scenes
- Knowledge driven Description Synthesis for Floor Plan Interpretation
- Side-Aware Boundary Localization for More Precise Object Detection
- Backdoor Attack through Frequency Domain
- Sequential Context Encoding for Duplicate Removal
- SemVLP: Vision-Language Pre-training by Aligning Semantics at Multiple Levels
- CurriculumNet: Weakly Supervised Learning from Large-Scale Web Images
- Tackling the Unannotated: Scene Graph Generation with Bias-Reduced Models
- Masked Vision-Language Transformer in Fashion
- Deep Perceptual Compression
- Towards Unified INT8 Training for Convolutional Neural Network
- Quantifying Legibility of Indoor Spaces Using Deep Convolutional Neural Networks: Case Studies in Train Stations
- SEED: Semantics Enhanced Encoder-Decoder Framework for Scene Text Recognition
- Language and Visual Entity Relationship Graph for Agent Navigation
- From Patches to Pictures (PaQ-2-PiQ): Mapping the Perceptual Space of Picture Quality
- Gated Feedback Refinement Network for Coarse-to-Fine Dense Semantic Image Labeling
- Co-localization with Category-Consistent Features and Geodesic Distance Propagation
- Learning to Recognize Actions on Objects in Egocentric Video with Attention Dictionaries
- Exploring Low-light Object Detection Techniques
- Alpha-Refine: Boosting Tracking Performance by Precise Bounding Box Estimation
- DFUC2020: Analysis Towards Diabetic Foot Ulcer Detection
- Certified Data Removal from Machine Learning Models
- Lightweight Convolutional Neural Network with Gaussian-based Grasping Representation for Robotic Grasping Detection
- A Scale Invariant Flatness Measure for Deep Network Minima
- PPR-FCN: Weakly Supervised Visual Relation Detection via Parallel Pairwise R-FCN
- Tackling the Challenges in Scene Graph Generation with Local-to-Global Interactions
- Pose-adaptive Hierarchical Attention Network for Facial Expression Recognition
- A Large-Scale Study on Unsupervised Spatiotemporal Representation Learning
- Implicit Feature Pyramid Network for Object Detection
- TAP: Text-Aware Pre-training for Text-VQA and Text-Caption
- Counterfactual Critic Multi-Agent Training for Scene Graph Generation
- A survey of Object Classification and Detection based on 2D/3D data
- GCNNMatch: Graph Convolutional Neural Networks for Multi-Object Tracking via Sinkhorn Normalization
- Baidu-UTS Submission to the EPIC-Kitchens Action Recognition Challenge 2019
- DRG: Dual Relation Graph for Human-Object Interaction Detection
- Domain-aware Visual Bias Eliminating for Generalized Zero-Shot Learning
- Detecting Text in Natural Image with Connectionist Text Proposal Network
- WSOD^2: Learning Bottom-up and Top-down Objectness Distillation for Weakly-supervised Object Detection
- 3D Interpreter Networks for Viewer-Centered Wireframe Modeling
- Learning Object Relation Graph and Tentative Policy for Visual Navigation
- Cost Volume Pyramid Based Depth Inference for Multi-View Stereo
- Task-Aware Variational Adversarial Active Learning
- Learning Layout and Style Reconfigurable GANs for Controllable Image Synthesis
- Deep Learning at 15PF: Supervised and Semi-Supervised Classification for Scientific Data
- CityFlow: A City-Scale Benchmark for Multi-Target Multi-Camera Vehicle Tracking and Re-Identification
- Conservation AI: Live Stream Analysis for the Detection of Endangered Species Using Convolutional Neural Networks and Drone Technology
- Empirical Upper Bound in Object Detection and More
- Consistent Optimization for Single-Shot Object Detection
- Neural Person Search Machines
- ReenactGAN: Learning to Reenact Faces via Boundary Transfer
- Classification of Findings with Localized Lesions in Fundoscopic Images using a Regionally Guided CNN
- Improving Image Captioning by Leveraging Intra- and Inter-layer Global Representation in Transformer Network
- Quasi-Dense Similarity Learning for Multiple Object Tracking
- Cross-domain Detection via Graph-induced Prototype Alignment
- Tracking Objects as Points
- Learning to decompose for object detection and instance segmentation
- Self-supervised learning for autonomous vehicles perception: A conciliation between analytical and learning methods
- ResNet or DenseNet? Introducing Dense Shortcuts to ResNet
- Dense Label Encoding for Boundary Discontinuity Free Rotation Detection
- Stories for Images-in-Sequence by using Visual and Narrative Components
- Contrastive Visual-Linguistic Pretraining
- C2FNAS: Coarse-to-Fine Neural Architecture Search for 3D Medical Image Segmentation
- Visual Parser: Representing Part-whole Hierarchies with Transformers
- Entropy-Enhanced Multimodal Attention Model for Scene-Aware Dialogue Generation
- SegVoxelNet: Exploring Semantic Context and Depth-aware Features for 3D Vehicle Detection from Point Cloud
- A Delay Metric for Video Object Detection: What Average Precision Fails to Tell
- Transformer-based Context Condensation for Boosting Feature Pyramids in Object Detection
- Target Detection, Tracking and Avoidance System for Low-cost UAVs using AI-Based Approaches
- Joint Detection and Tracking in Videos with Identification Features
- One-Shot Object Detection without Fine-Tuning
- Bridging Knowledge Graphs to Generate Scene Graphs
- In Defense of Grid Features for Visual Question Answering
- UPSNet: A Unified Panoptic Segmentation Network
- Joint Iris Segmentation and Localization Using Deep Multi-task Learning Framework
- Learning to Collocate Neural Modules for Image Captioning
- Look More Than Once: An Accurate Detector for Text of Arbitrary Shapes
- Regional Homogeneity: Towards Learning Transferable Universal Adversarial Perturbations Against Defenses
- Learning Salient Boundary Feature for Anchor-free Temporal Action Localization
- Bridge Damage Detection using a Single-Stage Detector and Field Inspection Images
- Dual Refinement Feature Pyramid Networks for Object Detection
- The Next Big Thing(s) in Unsupervised Machine Learning: Five Lessons from Infant Learning
- Hierarchical LSTMs with Adaptive Attention for Visual Captioning
- QPIC: Query-Based Pairwise Human-Object Interaction Detection with Image-Wide Contextual Information
- PIoU Loss: Towards Accurate Oriented Object Detection in Complex Environments
- Iterative Answer Prediction with Pointer-Augmented Multimodal Transformers for TextVQA
- TDAPNet: Prototype Network with Recurrent Top-Down Attention for Robust Object Classification under Partial Occlusion
- SPM-Tracker: Series-Parallel Matching for Real-Time Visual Object Tracking
- An interpretable classifier for high-resolution breast cancer screening images utilizing weakly supervised localization
- Action Genome: Actions as Composition of Spatio-temporal Scene Graphs
- SIENet: Spatial Information Enhancement Network for 3D Object Detection from Point Cloud
- Comparison Network for One-Shot Conditional Object Detection
- Person-in-WiFi: Fine-grained Person Perception using WiFi
- Siamese Box Adaptive Network for Visual Tracking
- Boundary Proposal Network for Two-Stage Natural Language Video Localization
- Large Scale Visual Food Recognition
- DeepWriter: A Multi-Stream Deep CNN for Text-independent Writer Identification
- Detailed 2D-3D Joint Representation for Human-Object Interaction
- Deep Optics for Monocular Depth Estimation and 3D Object Detection
- Multimodal Contrastive Training for Visual Representation Learning
- Temporal Recurrent Networks for Online Action Detection
- Revisiting Image-Language Networks for Open-ended Phrase Detection
- Two-Level Residual Distillation based Triple Network for Incremental Object Detection
- Cooperating RPN's Improve Few-Shot Object Detection
- Human-Object Interaction Prediction in Videos through Gaze Following
- Optimizing the Trade-off between Single-Stage and Two-Stage Object Detectors using Image Difficulty Prediction
- DeFRCN: Decoupled Faster R-CNN for Few-Shot Object Detection
- Universally Slimmable Networks and Improved Training Techniques
- From Deterministic to Generative: Multi-Modal Stochastic RNNs for Video Captioning
- A Deep Cascade of Convolutional Neural Networks for Dynamic MR Image Reconstruction
- Object-Based Visual Camera Pose Estimation From Ellipsoidal Model and 3D-Aware Ellipse Prediction
- Self-supervised Pre-training with Hard Examples Improves Visual Representations
- Object-aware Contrastive Learning for Debiased Scene Representation
- Saliency-Guided Attention Network for Image-Sentence Matching
- MotorEase: Automated Detection of Motor Impairment Accessibility Issues in Mobile App UIs
- Keep it Simple: Image Statistics Matching for Domain Adaptation
- Artificial and beneficial -- Exploiting artificial images for aerial vehicle detection
- GAN: Complementary Fashion Item Recommendation
- SEIGAN: Towards Compositional Image Generation by Simultaneously Learning to Segment, Enhance, and Inpaint
- Sperm Detection and Tracking in Phase-Contrast Microscopy Image Sequences using Deep Learning and Modified CSR-DCF
- TOOD: Task-aligned One-stage Object Detection
- Interact as You Intend: Intention-Driven Human-Object Interaction Detection
- Imperceptible Adversarial Examples by Spatial Chroma-Shift
- Ocean: Object-aware Anchor-free Tracking
- CentripetalText: An Efficient Text Instance Representation for Scene Text Detection
- MOTR: End-to-End Multiple-Object Tracking with Transformer
- Accurate Monocular Object Detection via Color-Embedded 3D Reconstruction for Autonomous Driving
- Structured Knowledge Distillation for Dense Prediction
- PanopticFusion: Online Volumetric Semantic Mapping at the Level of Stuff and Things
- Spherical Kernel for Efficient Graph Convolution on 3D Point Clouds
- Self-distillation with Batch Knowledge Ensembling Improves ImageNet Classification
- Signet Ring Cell Detection With a Semi-supervised Learning Framework
- Cross-dataset Training for Class Increasing Object Detection
- Improving Multispectral Pedestrian Detection by Addressing Modality Imbalance Problems
- Online Object-Oriented Semantic Mapping and Map Updating
- Multi-Scale Attention Network for Crowd Counting
- Stabilized Medical Image Attacks
- Post-Training Piecewise Linear Quantization for Deep Neural Networks
- Target Driven Instance Detection
- Cross-view Relation Networks for Mammogram Mass Detection
- ACDnet: An action detection network for real-time edge computing based on flow-guided feature approximation and memory aggregation
- Visual Concepts and Compositional Voting
- HAMBox: Delving into Online High-quality Anchors Mining for Detecting Outer Faces
- VIVO: Visual Vocabulary Pre-Training for Novel Object Captioning
- Loss Function Discovery for Object Detection via Convergence-Simulation Driven Search
- A Serverless Cloud-Fog Platform for DNN-Based Video Analytics with Incremental Learning
- Co-Mixup: Saliency Guided Joint Mixup with Supermodular Diversity
- RoIMix: Proposal-Fusion among Multiple Images for Underwater Object Detection
- Kernel Transformer Networks for Compact Spherical Convolution
- Dense RepPoints: Representing Visual Objects with Dense Point Sets
- Learning in the Frequency Domain
- RDSNet: A New Deep Architecture for Reciprocal Object Detection and Instance Segmentation
- Structured Multimodal Attentions for TextVQA
- Adaptive Object Detection Using Adjacency and Zoom Prediction
- Structured Binary Neural Networks for Accurate Image Classification and Semantic Segmentation
- Self-Knowledge Distillation with Progressive Refinement of Targets
- Learning Temporal Pose Estimation from Sparsely-Labeled Videos
- CTAP: Complementary Temporal Action Proposal Generation
- Query-guided End-to-End Person Search
- Point-Set Anchors for Object Detection, Instance Segmentation and Pose Estimation
- Align Deep Features for Oriented Object Detection
- Spatio-Temporal Action Detection with Cascade Proposal and Location Anticipation
- Deep Reasoning with Knowledge Graph for Social Relationship Understanding
- Objects in Semantic Topology
- BorderDet: Border Feature for Dense Object Detection
- Real-Time Seamless Single Shot 6D Object Pose Prediction
- Deeply Explain CNN via Hierarchical Decomposition
- CenterFace: Joint Face Detection and Alignment Using Face as Point
- Arbitrary-Shaped Text Detection withAdaptive Text Region Representation
- MoVie: Revisiting Modulated Convolutions for Visual Counting and Beyond
- MVOR: A Multi-view RGB-D Operating Room Dataset for 2D and 3D Human Pose Estimation
- Towards Noise-resistant Object Detection with Noisy Annotations
- A Comprehensive Approach for UAV Small Object Detection with Simulation-based Transfer Learning and Adaptive Fusion
- RetinaTrack: Online Single Stage Joint Detection and Tracking
- GENESIS-V2: Inferring Unordered Object Representations without Iterative Refinement
- Shape or Texture: Understanding Discriminative Features in CNNs
- Deep traffic light detection by overlaying synthetic context on arbitrary natural images
- ABN: Agent-Aware Boundary Networks for Temporal Action Proposal Generation
- Addressing the Cold-Start Problem in Outfit Recommendation Using Visual Preference Modelling
- Efficient Folded Attention for 3D Medical Image Reconstruction and Segmentation
- Rethinking Re-Sampling in Imbalanced Semi-Supervised Learning
- Hit-Detector: Hierarchical Trinity Architecture Search for Object Detection
- Contextual Multi-Scale Region Convolutional 3D Network for Activity Detection
- MoBiNet: A Mobile Binary Network for Image Classification
- Probabilistic Semantic Retrieval for Surveillance Videos with Activity Graphs
- Episode-based Prototype Generating Network for Zero-Shot Learning
- SimpleDet: A Simple and Versatile Distributed Framework for Object Detection and Instance Recognition
- Tracking by Instance Detection: A Meta-Learning Approach
- Cascaded Structure Tensor Framework for Robust Identification of Heavily Occluded Baggage Items from X-ray Scans
- Deep learning-based prediction of response to HER2-targeted neoadjuvant chemotherapy from pre-treatment dynamic breast MRI: A multi-institutional validation study
- Disp R-CNN: Stereo 3D Object Detection via Shape Prior Guided Instance Disparity Estimation
- BioDrone: A Bionic Drone-based Single Object Tracking Benchmark for Robust Vision
- Benchmarking Deep Trackers on Aerial Videos
- MMOCR: A Comprehensive Toolbox for Text Detection, Recognition and Understanding
- Segmentation is All You Need
- Looking in the Right place for Anomalies: Explainable AI through Automatic Location Learning
- Recurrent Autoregressive Networks for Online Multi-Object Tracking
- Robust Deep Neural Object Detection and Segmentation for Automotive Driving Scenario with Compressed Image Data
- A Cost-Effective Person-Following System for Assistive Unmanned Vehicles with Deep Learning at the Edge
- BoLTVOS: Box-Level Tracking for Video Object Segmentation
- Linguistically-aware Attention for Reducing the Semantic-Gap in Vision-Language Tasks
- Attention Neural Network for Trash Detection on Water Channels
- Dynamic Relevance Learning for Few-Shot Object Detection
- Deep Watershed Detector for Music Object Recognition
- Motion-Appearance Interactive Encoding for Object Segmentation in Unconstrained Videos
- Clustered Object Detection in Aerial Images
- DASS: Differentiable Architecture Search for Sparse neural networks
- Bipartite Graph Network with Adaptive Message Passing for Unbiased Scene Graph Generation
- Accurate RGB-D Salient Object Detection via Collaborative Learning
- Comprehensive Image Captioning via Scene Graph Decomposition
- Exploring the Design Space of Deep Convolutional Neural Networks at Large Scale
- Detecting Human-Object Interactions with Action Co-occurrence Priors
- Universal Semi-Supervised Semantic Segmentation
- Transfer Learning for Instance Segmentation of Waste Bottles using Mask R-CNN Algorithm
- SM-NAS: Structural-to-Modular Neural Architecture Search for Object Detection
- WebVision Challenge: Visual Learning and Understanding With Web Data
- Self-supervised 6D Object Pose Estimation for Robot Manipulation
- Deep GrabCut for Object Selection
- OVANet: One-vs-All Network for Universal Domain Adaptation
- Actor-Centric Relation Network
- A Comparative Review of Recent Few-Shot Object Detection Algorithms
- COIN: A Large-scale Dataset for Comprehensive Instructional Video Analysis
- Learning Efficient Detector with Semi-supervised Adaptive Distillation
- IoU-aware Single-stage Object Detector for Accurate Localization
- VisualGPT: Data-efficient Adaptation of Pretrained Language Models for Image Captioning
- Counting and Segmenting Sorghum Heads
- Every Pixel Matters: Center-aware Feature Alignment for Domain Adaptive Object Detector
- Wasserstein Distance Based Domain Adaptation for Object Detection
- Move to See Better: Self-Improving Embodied Object Detection
- End-to-End Object Detection with Fully Convolutional Network
- Memory Warps for Learning Long-Term Online Video Representations
- Person Recognition in Personal Photo Collections
- Human-centric Spatio-Temporal Video Grounding With Visual Transformers
- Generalization Bounds for Convolutional Neural Networks
- ASFD: Automatic and Scalable Face Detector
- Feature Intertwiner for Object Detection
- Two-Stream Consensus Network for Weakly-Supervised Temporal Action Localization
- Rethinking Channel Dimensions for Efficient Model Design
- Structure-Aware Completion of Photogrammetric Meshes in Urban Road Environment
- Making an Invisibility Cloak: Real World Adversarial Attacks on Object Detectors
- Video Instance Segmentation
- Label-PEnet: Sequential Label Propagation and Enhancement Networks for Weakly Supervised Instance Segmentation
- When Healthcare Meets Off-the-Shelf WiFi: A Non-Wearable and Low-Costs Approach for In-Home Monitoring
- SWIPENET: Object detection in noisy underwater images
- Visual Object Recognition in Indoor Environments Using Topologically Persistent Features
- Hateful Memes Detection via Complementary Visual and Linguistic Networks
- Dynamic Mini-batch SGD for Elastic Distributed Training: Learning in the Limbo of Resources
- HRCenterNet: An Anchorless Approach to Chinese Character Segmentation in Historical Documents
- A Configurable BNN ASIC using a Network of Programmable Threshold Logic Standard Cells
- AdderNet: Do We Really Need Multiplications in Deep Learning?
- Dynamic Refinement Network for Oriented and Densely Packed Object Detection
- Cross-Modal Graph with Meta Concepts for Video Captioning
- Adversarial Learning of Structure-Aware Fully Convolutional Networks for Landmark Localization
- Pipelines for Procedural Information Extraction from Scientific Literature: Towards Recipes using Machine Learning and Data Science
- An Analysis of Scale Invariance in Object Detection - SNIP
- SceneGen: Generative Contextual Scene Augmentation using Scene Graph Priors
- Pose2Seg: Detection Free Human Instance Segmentation
- A Fast Face Detection Method via Convolutional Neural Network
- MovieNet: A Holistic Dataset for Movie Understanding
- Spatial Semantic Embedding Network: Fast 3D Instance Segmentation with Deep Metric Learning
- HPC AI500: The Methodology, Tools, Roofline Performance Models, and Metrics for Benchmarking HPC AI Systems
- IA-MOT: Instance-Aware Multi-Object Tracking with Motion Consistency
- MaskFace: multi-task face and landmark detector
- Train in Germany, Test in The USA: Making 3D Object Detectors Generalize
- Visual Relationship Detection using Scene Graphs: A Survey
- WQT and DG-YOLO: towards domain generalization in underwater object detection
- Modeling Visual Context is Key to Augmenting Object Detection Datasets
- TOG: Targeted Adversarial Objectness Gradient Attacks on Real-time Object Detection Systems
- SG-Net: Spatial Granularity Network for One-Stage Video Instance Segmentation
- Convolutional Recurrent Predictor: Implicit Representation for Multi-target Filtering and Tracking
- KeepAugment: A Simple Information-Preserving Data Augmentation Approach
- Slender Object Detection: Diagnoses and Improvements
- Learning Transferable Adversarial Examples via Ghost Networks
- SPLAT: Semantic Pixel-Level Adaptation Transforms for Detection
- Pose Estimation for Non-Cooperative Rendezvous Using Neural Networks
- AM-LFS: AutoML for Loss Function Search
- Hallucinating IDT Descriptors and I3D Optical Flow Features for Action Recognition with CNNs
- NeRF in detail: Learning to sample for view synthesis
- Amodal Instance Segmentation
- Attribute Aware Pooling for Pedestrian Attribute Recognition
- Look at What I'm Doing: Self-Supervised Spatial Grounding of Narrations in Instructional Videos
- Instance Segmentation with Point Supervision
- AD-Det: Boosting Object Detection in UAV Images with Focused Small Objects and Balanced Tail Classes
- Skew Class-balanced Re-weighting for Unbiased Scene Graph Generation
- Synthetic Document Generator for Annotation-free Layout Recognition
- Are You Looking? Grounding to Multiple Modalities in Vision-and-Language Navigation
- Online Object Representations with Contrastive Learning
- PolyTransform: Deep Polygon Transformer for Instance Segmentation
- Fashion Retrieval via Graph Reasoning Networks on a Similarity Pyramid
- Class-Incremental Few-Shot Object Detection
- PASS3D: Precise and Accelerated Semantic Segmentation for 3D Point Cloud
- Machine-learning based methodologies for 3d x-ray measurement, characterization and optimization for buried structures in advanced ic packages
- Scan2Plan: Efficient Floorplan Generation from 3D Scans of Indoor Scenes
- Accurate 6D Object Pose Estimation by Pose Conditioned Mesh Reconstruction
- Track to Detect and Segment: An Online Multi-Object Tracker
- Learning Deep Representations for Semantic Image Parsing: a Comprehensive Overview
- Let's measure run time! Extending the IR replicability infrastructure to include performance aspects
- Semi-supervised Active Learning for Instance Segmentation via Scoring Predictions
- MimicDet: Bridging the Gap Between One-Stage and Two-Stage Object Detection
- Removing the Background by Adding the Background: Towards Background Robust Self-supervised Video Representation Learning
- Soft Anchor-Point Object Detection
- Making History Matter: History-Advantage Sequence Training for Visual Dialog
- Breast Mass Detection with Faster R-CNN: On the Feasibility of Learning from Noisy Annotations
- ClipCap: CLIP Prefix for Image Captioning
- Concurrent Segmentation and Object Detection CNNs for Aircraft Detection and Identification in Satellite Images
- FBNetV5: Neural Architecture Search for Multiple Tasks in One Run
- Visual Compositional Learning for Human-Object Interaction Detection
- Feature Pyramid Grids
- More Grounded Image Captioning by Distilling Image-Text Matching Model
- Selecting Relevant Features from a Multi-domain Representation for Few-shot Classification
- Semi-supervised Open-World Object Detection
- Improving Description-based Person Re-identification by Multi-granularity Image-text Alignments
- VoxelPose: Towards Multi-Camera 3D Human Pose Estimation in Wild Environment
- Fully Convolutional Networks for Panoptic Segmentation
- Point Cloud Instance Segmentation using Probabilistic Embeddings
- The NAO Backpack: An Open-hardware Add-on for Fast Software Development with the NAO Robot
- Emotion Recognition From Gait Analyses: Current Research and Future Directions
- Decoupling the Role of Data, Attention, and Losses in Multimodal Transformers
- Beef Cattle Instance Segmentation Using Fully Convolutional Neural Network
- Convolutional Auto-encoding of Sentence Topics for Image Paragraph Generation
- Range Conditioned Dilated Convolutions for Scale Invariant 3D Object Detection
- Weakly Supervised Complementary Parts Models for Fine-Grained Image Classification from the Bottom Up
- Image-based localization using LSTMs for structured feature correlation
- A Dataset for Provident Vehicle Detection at Night
- Learning Feature-to-Feature Translator by Alternating Back-Propagation for Generative Zero-Shot Learning
- A 2D laser rangefinder scans dataset of standard EUR pallets
- Depth-conditioned Dynamic Message Propagation for Monocular 3D Object Detection
- Fast, Diverse and Accurate Image Captioning Guided By Part-of-Speech
- Appearance-Preserving 3D Convolution for Video-based Person Re-identification
- A Deep Learning-based Framework for the Detection of Schools of Herring in Echograms
- Learning a Layout Transfer Network for Context Aware Object Detection
- Naive-Student: Leveraging Semi-Supervised Learning in Video Sequences for Urban Scene Segmentation
- Pose Neural Fabrics Search
- Inter-Image Communication for Weakly Supervised Localization
- TRIE: End-to-End Text Reading and Information Extraction for Document Understanding
- 6-PACK: Category-level 6D Pose Tracker with Anchor-Based Keypoints
- Towards Linking the Lakh and IMSLP Datasets
- Progressive Localization Networks for Language-based Moment Localization
- Online Multiple Pedestrians Tracking using Deep Temporal Appearance Matching Association
- Recurrence along Depth: Deep Convolutional Neural Networks with Recurrent Layer Aggregation
- Dual Attention Networks for Visual Reference Resolution in Visual Dialog
- Cooper: Cooperative Perception for Connected Autonomous Vehicles based on 3D Point Clouds
- AEI: Actors-Environment Interaction with Adaptive Attention for Temporal Action Proposals Generation
- On Learning Vehicle Detection in Satellite Video
- Analysis of diversity-accuracy tradeoff in image captioning
- CircleNet: Reciprocating Feature Adaptation for Robust Pedestrian Detection
- Rethinking Differentiable Search for Mixed-Precision Neural Networks
- Towards Compact ConvNets via Structure-Sparsity Regularized Filter Pruning
- Temporal HeartNet: Towards Human-Level Automatic Analysis of Fetal Cardiac Screening Video
- Exploiting long-term temporal dynamics for video captioning
- PaStaNet: Toward Human Activity Knowledge Engine
- An Analysis of Deep Object Detectors For Diver Detection
- The Garden of Forking Paths: Towards Multi-Future Trajectory Prediction
- Attentive Relational Networks for Mapping Images to Scene Graphs
- Guided Attention Network for Object Detection and Counting on Drones
- A Free Lunch for Unsupervised Domain Adaptive Object Detection without Source Data
- StartNet: Online Detection of Action Start in Untrimmed Videos
- VIPriors 1: Visual Inductive Priors for Data-Efficient Deep Learning Challenges
- Tracking objects using 3D object proposals
- Meta Module Network for Compositional Visual Reasoning
- Joint Contrastive Learning with Infinite Possibilities
- ScanSSD: Scanning Single Shot Detector for Mathematical Formulas in PDF Document Images
- AI Matrix: A Deep Learning Benchmark for Alibaba Data Centers
- A Software Architecture for Autonomous Vehicles: Team LRM-B Entry in the First CARLA Autonomous Driving Challenge
- Refinements in Motion and Appearance for Online Multi-Object Tracking
- CGaP: Continuous Growth and Pruning for Efficient Deep Learning
- HRank: Filter Pruning using High-Rank Feature Map
- Evaluating Generalization Ability of Convolutional Neural Networks and Capsule Networks for Image Classification via Top-2 Classification
- Detecting, Localising and Classifying Polyps from Colonoscopy Videos using Deep Learning
- Jointly Cross- and Self-Modal Graph Attention Network for Query-Based Moment Localization
- Reformulating HOI Detection as Adaptive Set Prediction
- Geometry Normalization Networks for Accurate Scene Text Detection
- Scaling-Translation-Equivariant Networks with Decomposed Convolutional Filters
- HiFT: Hierarchical Feature Transformer for Aerial Tracking
- Autonomous Marine Sampling Enhanced by Strategically Deployed Drifters in Marine Flow Fields
- Deep Frequent Spatial Temporal Learning for Face Anti-Spoofing
- DeNet: Scalable Real-time Object Detection with Directed Sparse Sampling
- Fast Video Shot Transition Localization with Deep Structured Models
- Abstractive Text Classification Using Sequence-to-convolution Neural Networks
- Hallucination Improves Few-Shot Object Detection
- Detecting Noteheads in Handwritten Scores with ConvNets and Bounding Box Regression
- Self-Adversarial Disentangling for Specific Domain Adaptation
- How benign is benign overfitting?
- The Role of Context Selection in Object Detection
- Single-Shot 3D Detection of Vehicles from Monocular RGB Images via Geometry Constrained Keypoints in Real-Time
- Fusion of Head and Full-Body Detectors for Multi-Object Tracking
- Attention Allocation Aid for Visual Search
- Generative Partition Networks for Multi-Person Pose Estimation
- Training a Fast Object Detector for LiDAR Range Images Using Labeled Data from Sensors with Higher Resolution
- Siamese Keypoint Prediction Network for Visual Object Tracking
- Cross-modal Scene Graph Matching for Relationship-aware Image-Text Retrieval
- Deep Affinity Net: Instance Segmentation via Affinity
- SemiNLL: A Framework of Noisy-Label Learning by Semi-Supervised Learning
- Learning Video Representations from Correspondence Proposals
- VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering
- Single Pixel Reconstruction for One-stage Instance Segmentation
- Automated Steel Bar Counting and Center Localization with Convolutional Neural Networks
- Instance Segmentation in 3D Scenes using Semantic Superpoint Tree Networks
- Multiview Detection with Feature Perspective Transformation
- An End-to-End Network for Co-Saliency Detection in One Single Image
- Customizing Student Networks From Heterogeneous Teachers via Adaptive Knowledge Amalgamation
- Affinity LCFCN: Learning to Segment Fish with Weak Supervision
- CO2: Consistent Contrast for Unsupervised Visual Representation Learning
- MFPN: A Novel Mixture Feature Pyramid Network of Multiple Architectures for Object Detection
- Down to the Last Detail: Virtual Try-on with Detail Carving
- Double Anchor R-CNN for Human Detection in a Crowd
- Matching Images and Text with Multi-modal Tensor Fusion and Re-ranking
- Spatiotemporal Relationship Reasoning for Pedestrian Intent Prediction
- X-Linear Attention Networks for Image Captioning
- Spatially Consistent Representation Learning
- Real-Time Shape Tracking of Facial Landmarks
- Co-Separating Sounds of Visual Objects
- PanoNet: Real-time Panoptic Segmentation through Position-Sensitive Feature Embedding
- An End-to-End Approach for Recognition of Modern and Historical Handwritten Numeral Strings
- OmDet: Large-scale vision-language multi-dataset pre-training with multimodal detection network
- Fashion Captioning: Towards Generating Accurate Descriptions with Semantic Rewards
- Sex Trafficking Detection with Ordinal Regression Neural Networks
- Deep neural networks approach to microbial colony detection -- a comparative analysis
- 3D-Aware Ellipse Prediction for Object-Based Camera Pose Estimation
- Lightweight Pyramid Networks for Image Deraining
- Sparse Weight Activation Training
- Prediction-Tracking-Segmentation
- MMFT-BERT: Multimodal Fusion Transformer with BERT Encodings for Visual Question Answering
- Rethinking Self-Supervised Learning: Small is Beautiful
- Attribute And-Or Grammar for Joint Parsing of Human Attributes, Part and Pose
- VQA-LOL: Visual Question Answering under the Lens of Logic
- Learning Region Features for Object Detection
- NASGEM: Neural Architecture Search via Graph Embedding Method
- Open Set Recognition with Conditional Probabilistic Generative Models
- Advances in Deep Learning for Hyperspectral Image Analysis--Addressing Challenges Arising in Practical Imaging Scenarios
- Sketching Image Gist: Human-Mimetic Hierarchical Scene Graph Generation
- End-to-End Wireframe Parsing
- A Hybrid Approach and Unified Framework for Bibliographic Reference Extraction
- Sitatapatra: Blocking the Transfer of Adversarial Samples
- Learning from Lexical Perturbations for Consistent Visual Question Answering
- OMNIA Faster R-CNN: Detection in the wild through dataset merging and soft distillation
- A Survey on Deep Domain Adaptation and Tiny Object Detection Challenges, Techniques and Datasets
- Learning Where to Focus for Efficient Video Object Detection
- 3DIoUMatch: Leveraging IoU Prediction for Semi-Supervised 3D Object Detection
- MSG-Transformer: Exchanging Local Spatial Information by Manipulating Messenger Tokens
- A Unified Benchmark for the Unknown Detection Capability of Deep Neural Networks
- Object-Centric Diagnosis of Visual Reasoning
- Toward Automatic Threat Recognition for Airport X-ray Baggage Screening with Deep Convolutional Object Detection
- Mask R-CNN with Pyramid Attention Network for Scene Text Detection
- Localizing Actions from Video Labels and Pseudo-Annotations
- SpatialFlow: Bridging All Tasks for Panoptic Segmentation
- AutoCaption: Image Captioning with Neural Architecture Search
- VIOLIN: A Large-Scale Dataset for Video-and-Language Inference
- MSNet: A Multilevel Instance Segmentation Network for Natural Disaster Damage Assessment in Aerial Videos
- MnasFPN: Learning Latency-aware Pyramid Architecture for Object Detection on Mobile Devices
- Table Structure Recognition using Top-Down and Bottom-Up Cues
- Plumber: Diagnosing and Removing Performance Bottlenecks in Machine Learning Data Pipelines
- Large-Scale Classification of Structured Objects using a CRF with Deep Class Embedding
- Harvesting, Detecting, and Characterizing Liver Lesions from Large-scale Multi-phase CT Data via Deep Dynamic Texture Learning
- SMOT: Single-Shot Multi Object Tracking
- Towards Phytoplankton Parasite Detection Using Autoencoders
- BinaryRelax: A Relaxation Approach For Training Deep Neural Networks With Quantized Weights
- Road Damage Detection and Classification with Detectron2 and Faster R-CNN
- Revisiting the Sibling Head in Object Detector
- Localization in the Crowd with Topological Constraints
- Predicting Electricity Consumption using Deep Recurrent Neural Networks
- MTL-NAS: Task-Agnostic Neural Architecture Search towards General-Purpose Multi-Task Learning
- MotionNet: Joint Perception and Motion Prediction for Autonomous Driving Based on Bird's Eye View Maps
- Probing Inter-modality: Visual Parsing with Self-Attention for Vision-Language Pre-training
- Combating Uncertainty with Novel Losses for Automatic Left Atrium Segmentation
- Finding and Visualizing Weaknesses of Deep Reinforcement Learning Agents
- Progressive Correspondence Pruning by Consensus Learning
- Dual-Awareness Attention for Few-Shot Object Detection
- Stacked Cross-modal Feature Consolidation Attention Networks for Image Captioning
- DR Loss: Improving Object Detection by Distributional Ranking
- Attention: A Big Surprise for Cross-Domain Person Re-Identification
- EXSCLAIM! -- An automated pipeline for the construction of labeled materials imaging datasets from literature
- Learning Pairwise Relationship for Multi-object Detection in Crowded Scenes
- VisDA-2021 Competition Universal Domain Adaptation to Improve Performance on Out-of-Distribution Data
- Point in, Box out: Beyond Counting Persons in Crowds
- Fine-Grained Image Analysis with Deep Learning: A Survey
- A Gap-Based Framework for Chinese Word Segmentation via Very Deep Convolutional Networks
- Diversifying Sample Generation for Accurate Data-Free Quantization
- NoduleNet: Decoupled False Positive Reductionfor Pulmonary Nodule Detection and Segmentation
- Rethinking Classification and Localization for Cascade R-CNN
- LogoDet-3K: A Large-Scale Image Dataset for Logo Detection
- Rethinking Pseudo-LiDAR Representation
- Pose-based Modular Network for Human-Object Interaction Detection
- Driver Behavior Analysis Using Lane Departure Detection Under Challenging Conditions
- Learning to Combine: Knowledge Aggregation for Multi-Source Domain Adaptation
- Unsupervised Hard Example Mining from Videos for Improved Object Detection
- Co-training for Deep Object Detection: Comparing Single-modal and Multi-modal Approaches
- GPS-Net: Graph Property Sensing Network for Scene Graph Generation
- STINet: Spatio-Temporal-Interactive Network for Pedestrian Detection and Trajectory Prediction
- Improving Calibration for Long-Tailed Recognition
- Matching Visual Features to Hierarchical Semantic Topics for Image Paragraph Captioning
- Mask TextSpotter v3: Segmentation Proposal Network for Robust Scene Text Spotting
- End-to-End Pseudo-LiDAR for Image-Based 3D Object Detection
- Discovering Visual Patterns in Art Collections with Spatially-consistent Feature Learning
- Weakly-Supervised Spatio-Temporally Grounding Natural Sentence in Video
- DirectShape: Direct Photometric Alignment of Shape Priors for Visual Vehicle Pose and Shape Estimation
- Deep Convolutional Correlation Iterative Particle Filter for Visual Tracking
- Proposal-based Few-shot Sound Event Detection for Speech and Environmental Sounds with Perceivers
- Inference, Learning and Attention Mechanisms that Exploit and Preserve Sparsity in Convolutional Networks
- IAN: The Individual Aggregation Network for Person Search
- DocVQA: A Dataset for VQA on Document Images
- Forest R-CNN: Large-Vocabulary Long-Tailed Object Detection and Instance Segmentation
- LLA: Loss-aware Label Assignment for Dense Pedestrian Detection
- Convolutional Character Networks
- Evaluation of Momentum Diverse Input Iterative Fast Gradient Sign Method (M-DI2-FGSM) Based Attack Method on MCS 2018 Adversarial Attacks on Black Box Face Recognition System
- Learning from Multiple Datasets with Heterogeneous and Partial Labels for Universal Lesion Detection in CT
- Automatic Information Extraction from Piping and Instrumentation Diagrams
- Self-Supervised Learning of Depth and Ego-motion with Differentiable Bundle Adjustment
- G2L-Net: Global to Local Network for Real-time 6D Pose Estimation with Embedding Vector Features
- Finding the Evidence: Localization-aware Answer Prediction for Text Visual Question Answering
- An Automatic System for Unconstrained Video-Based Face Recognition
- RarePlanes: Synthetic Data Takes Flight
- Object Detection in Video with Spatial-temporal Context Aggregation
- Learning a Disentangled Embedding for Monocular 3D Shape Retrieval and Pose Estimation
- S2DNAS:Transforming Static CNN Model for Dynamic Inference via Neural Architecture Search
- Adaptive Label Smoothing
- Improving Weakly Supervised Visual Grounding by Contrastive Knowledge Distillation
- Question Type Guided Attention in Visual Question Answering
- ALFA: Agglomerative Late Fusion Algorithm for Object Detection
- Decision-based Universal Adversarial Attack
- CenterMask: single shot instance segmentation with point representation
- 3D for Free: Crossmodal Transfer Learning using HD Maps
- Semi-convolutional Operators for Instance Segmentation
- Learning Cross-modal Context Graph for Visual Grounding
- Gesture Recognition for Initiating Human-to-Robot Handovers
- In the Eye of the Beholder: Gaze and Actions in First Person Video
- Pyramid Multi-view Stereo Net with Self-adaptive View Aggregation
- Local Metrics for Multi-Object Tracking
- Depth-wise Decomposition for Accelerating Separable Convolutions in Efficient Convolutional Neural Networks
- Unbiased Scene Graph Generation via Rich and Fair Semantic Extraction
- ReFormer: The Relational Transformer for Image Captioning
- Scheduled Sampling in Vision-Language Pretraining with Decoupled Encoder-Decoder Network
- Improve Object Detection by Data Enhancement based on Generative Adversarial Nets
- Self-Supervised Learning for Semi-Supervised Temporal Action Proposal
- Adversarial Cross-Domain Action Recognition with Co-Attention
- Intention Recognition of Pedestrians and Cyclists by 2D Pose Estimation
- Unsupervised Multiple Person Tracking using AutoEncoder-Based Lifted Multicuts
- Open Domain Generalization with Domain-Augmented Meta-Learning
- Effective and Robust Detection of Adversarial Examples via Benford-Fourier Coefficients
- Multi-Label Generalized Zero Shot Learning for the Classification of Disease in Chest Radiographs
- S3VAE: Self-Supervised Sequential VAE for Representation Disentanglement and Data Generation
- TracKlinic: Diagnosis of Challenge Factors in Visual Tracking
- A Multi Camera Unsupervised Domain Adaptation Pipeline for Object Detection in Cultural Sites through Adversarial Learning and Self-Training
- Weakly Supervised Cell Instance Segmentation by Propagating from Detection Response
- Act, Perceive, and Plan in Belief Space for Robot Localization
- ObjectNet Dataset: Reanalysis and Correction
- Contextual Heterogeneous Graph Network for Human-Object Interaction Detection
- Team JL Solution to Google Landmark Recognition 2019
- Category-wise Attack: Transferable Adversarial Examples for Anchor Free Object Detection
- StrObe: Streaming Object Detection from LiDAR Packets
- PosMLP-Video: Spatial and Temporal Relative Position Encoding for Efficient Video Recognition
- Benchmarking the Robustness of Instance Segmentation Models
- Exploiting Kernel Sparsity and Entropy for Interpretable CNN Compression
- VLUC: An Empirical Benchmark for Video-Like Urban Computing on Citywide Crowd and Traffic Prediction
- Iterative Low-Rank Approximation for CNN Compression
- Low-Power Object Counting with Hierarchical Neural Networks
- Estimating and Evaluating Regression Predictive Uncertainty in Deep Object Detectors
- Multi-Task Multi-Sensor Fusion for 3D Object Detection
- DSFD: Dual Shot Face Detector
- Improving Pixel Embedding Learning through Intermediate Distance Regression Supervision for Instance Segmentation
- RfD-Net: Point Scene Understanding by Semantic Instance Reconstruction
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal Localization
- Exploiting Both Domain-specific and Invariant Knowledge via a Win-win Transformer for Unsupervised Domain Adaptation
- Real-time Mask Detection on Google Edge TPU
- LapNet : Automatic Balanced Loss and Optimal Assignment for Real-Time Dense Object Detection
- Two-Stream AMTnet for Action Detection
- Instance Localization for Self-supervised Detection Pretraining
- Cascaded Human-Object Interaction Recognition
- MM-FSOD: Meta and metric integrated few-shot object detection
- MiLeNAS: Efficient Neural Architecture Search via Mixed-Level Reformulation
- Rethinking Normalization and Elimination Singularity in Neural Networks
- Bottom-Up Temporal Action Localization with Mutual Regularization
- UAV Visual Teach and Repeat Using Only Semantic Object Features
- Learning Context Graph for Person Search
- Towards a Multi-modal, Multi-task Learning based Pre-training Framework for Document Representation Learning
- CircleNet: Anchor-free Detection with Circle Representation
- Trimmed Action Recognition, Dense-Captioning Events in Videos, and Spatio-temporal Action Localization with Focus on ActivityNet Challenge 2019
- A Deep Neuro-Fuzzy Network for Image Classification
- TrackNet: A Deep Learning Network for Tracking High-speed and Tiny Objects in Sports Applications
- Semantic keypoint-based pose estimation from single RGB frames
- Fruit Quantity and Quality Estimation using a Robotic Vision System
- CARAFE++: Unified Content-Aware ReAssembly of FEatures
- Image Transformation can make Neural Networks more robust against Adversarial Examples
- Deep Learning Based Computed Tomography Whys and Wherefores
- Deep Occlusion-Aware Instance Segmentation with Overlapping BiLayers
- Multi-Sensor 3D Object Box Refinement for Autonomous Driving
- Cross-media Structured Common Space for Multimedia Event Extraction
- Style Normalization and Restitution for Domain Generalization and Adaptation
- What Vision-Language Models `See' when they See Scenes
- Enhancing Object Detection in Adverse Conditions using Thermal Imaging
- Learning latent representations across multiple data domains using Lifelong VAEGAN
- BoundarySqueeze: Image Segmentation as Boundary Squeezing
- Regularized Two-Branch Proposal Networks for Weakly-Supervised Moment Retrieval in Videos
- E2E-VLP: End-to-End Vision-Language Pre-training Enhanced by Visual Learning
- Memorizing Comprehensively to Learn Adaptively: Unsupervised Cross-Domain Person Re-ID with Multi-level Memory
- LGPMA: Complicated Table Structure Recognition with Local and Global Pyramid Mask Alignment
- Multilayer Collaborative Low-Rank Coding Network for Robust Deep Subspace Discovery
- PIXOR: Real-time 3D Object Detection from Point Clouds
- A^2-Net: Molecular Structure Estimation from Cryo-EM Density Volumes
- Image Captioning Based on a Hierarchical Attention Mechanism and Policy Gradient Optimization
- ViP-DeepLab: Learning Visual Perception with Depth-aware Video Panoptic Segmentation
- Positional Encoding as Spatial Inductive Bias in GANs
- Single-shot Path Integrated Panoptic Segmentation
- Patch2Pix: Epipolar-Guided Pixel-Level Correspondences
- Malaria Detection and Classificaiton
- Accelerating Minibatch Stochastic Gradient Descent using Typicality Sampling
- Cooperative Holistic Scene Understanding: Unifying 3D Object, Layout, and Camera Pose Estimation
- Real-world Mapping of Gaze Fixations Using Instance Segmentation for Road Construction Safety Applications
- Rethinking Natural Adversarial Examples for Classification Models
- Multi-person Articulated Tracking with Spatial and Temporal Embeddings
- Automatic microscopic cell counting by use of unsupervised adversarial domain adaptation and supervised density regression
- Effective and Fast: A Novel Sequential Single Path Search for Mixed-Precision Quantization
- Pan-tilt-zoom SLAM for Sports Videos
- Real-Time Text Detection and Recognition
- An Overview Of 3D Object Detection
- TAN: Temporal Affine Network for Real-Time Left Ventricle Anatomical Structure Analysis Based on 2D Ultrasound Videos
- Disjoint Label Space Transfer Learning with Common Factorised Space
- On the Importance of Visual Context for Data Augmentation in Scene Understanding
- Training Deep Neural Networks with Joint Quantization and Pruning of Weights and Activations
- You Only Look One-level Feature
- hSDB-instrument: Instrument Localization Database for Laparoscopic and Robotic Surgeries
- Dense Relational Captioning: Triple-Stream Networks for Relationship-Based Captioning
- Robust Glare Detection: Review, Analysis, and Dataset Release
- LDC-Net: A Unified Framework for Localization, Detection and Counting in Dense Crowds
- Adversarial Attacks in a Multi-view Setting: An Empirical Study of the Adversarial Patches Inter-view Transferability
- Synergizing between Self-Training and Adversarial Learning for Domain Adaptive Object Detection
- Detecting Human-Object Interaction via Fabricated Compositional Learning
- OTA: Optimal Transport Assignment for Object Detection
- Spatio-temporal Person Retrieval via Natural Language Queries
- GrooMeD-NMS: Grouped Mathematically Differentiable NMS for Monocular 3D Object Detection
- Asymptotic Soft Filter Pruning for Deep Convolutional Neural Networks
- Dynamic Temporal Pyramid Network: A Closer Look at Multi-Scale Modeling for Activity Detection
- Elucidating image-to-set prediction: An analysis of models, losses and datasets
- RevealNet: Seeing Behind Objects in RGB-D Scans
- Localizing Unseen Activities in Video via Image Query
- Deep Rigid Instance Scene Flow
- Weakly-Supervised Spatio-Temporal Anomaly Detection in Surveillance Video
- LAMPRET: Layout-Aware Multimodal PreTraining for Document Understanding
- Spectro-Temporal RF Identification using Deep Learning
- Persuasive Faces: Generating Faces in Advertisements
- DeSTNet: Densely Fused Spatial Transformer Networks
- Adaptation Across Extreme Variations using Unlabeled Domain Bridges
- Joint Visual Grounding with Language Scene Graphs
- Graph-Structured Referring Expression Reasoning in The Wild
- Multiple Object Tracking with Mixture Density Networks for Trajectory Estimation
- Towards End-to-End Text Spotting in Natural Scenes
- The Problem of Fragmented Occlusion in Object Detection
- A Local-to-Global Approach to Multi-modal Movie Scene Segmentation
- The EPIC-KITCHENS Dataset: Collection, Challenges and Baselines
- Towards Compact CNNs via Collaborative Compression
- Unifying Training and Inference for Panoptic Segmentation
- Towards the Internet of Robotic Things: Analysis, Architecture, Components and Challenges
- Using Artificial Intelligence to Analyze Fashion Trends
- The Newspaper Navigator Dataset: Extracting And Analyzing Visual Content from 16 Million Historic Newspaper Pages in Chronicling America
- Rethinking the Route Towards Weakly Supervised Object Localization
- Video-Text Pre-training with Learned Regions
- Where Does It Exist: Spatio-Temporal Video Grounding for Multi-Form Sentences
- Actions as Moving Points
- Detecting Semantic Parts on Partially Occluded Objects
- Vision-based Robot Manipulation Learning via Human Demonstrations
- Universal Lesion Detection by Learning from Multiple Heterogeneously Labeled Datasets
- APRICOT: A Dataset of Physical Adversarial Attacks on Object Detection
- Acceleration of Actor-Critic Deep Reinforcement Learning for Visual Grasping in Clutter by State Representation Learning Based on Disentanglement of a Raw Input Image
- AutoScale: Learning to Scale for Crowd Counting and Localization
- Image Enhanced Rotation Prediction for Self-Supervised Learning
- Action Recognition for Depth Video using Multi-view Dynamic Images
- Learning Filter Basis for Convolutional Neural Network Compression
- Robust Attentive Deep Neural Network for Exposing GAN-generated Faces
- Anyone here? Smart embedded low-resolution omnidirectional video sensor to measure room occupancy
- Towards explainable artificial intelligence (XAI) for early anticipation of traffic accidents
- Unsupervised Object-Level Representation Learning from Scene Images
- VSR: A Unified Framework for Document Layout Analysis combining Vision, Semantics and Relations
- Scene Understanding for Autonomous Driving
- Curiosity-driven Reinforcement Learning for Diverse Visual Paragraph Generation
- Bring Your Own Codegen to Deep Learning Compiler
- Feature Pyramid Transformer
- Visual Diagnosis of Dermatological Disorders: Human and Machine Performance
- Unsupervised Deep Representation Learning for Real-Time Tracking
- Parts4Feature: Learning 3D Global Features from Generally Semantic Parts in Multiple Views
- Propagating Asymptotic-Estimated Gradients for Low Bitwidth Quantized Neural Networks
- Affordance Transfer Learning for Human-Object Interaction Detection
- Unsupervised Discovery of Object Landmarks as Structural Representations
- Recovering hard-to-find object instances by sampling context-based object proposals
- Unmanned Aerial Vehicle Visual Detection and Tracking using Deep Neural Networks: A Performance Benchmark
- ClickBAIT: Click-based Accelerated Incremental Training of Convolutional Neural Networks
- Gradually Applying Weakly Supervised and Active Learning for Mass Detection in Breast Ultrasound Images
- Robust RGB-based 6-DoF Pose Estimation without Real Pose Annotations
- Towards Practical Implementations of Person Re-Identification from Full Video Frames
- Modeling and Analysis of Energy Harvesting and Smart Grid-Powered Wireless Communication Networks: A Contemporary Survey
- ApproxDet: Content and Contention-Aware Approximate Object Detection for Mobiles
- Weakly Supervised Learning of Multi-Object 3D Scene Decompositions Using Deep Shape Priors
- Regularizing Neural Networks via Adversarial Model Perturbation
- Feature Selective Small Object Detection via Knowledge-based Recurrent Attentive Neural Network
- SGPN: Similarity Group Proposal Network for 3D Point Cloud Instance Segmentation
- CogTree: Cognition Tree Loss for Unbiased Scene Graph Generation
- Meta-learning algorithms for Few-Shot Computer Vision
- From Selective Deep Convolutional Features to Compact Binary Representations for Image Retrieval
- A Novel Automation-Assisted Cervical Cancer Reading Method Based on Convolutional Neural Network
- A Deep Learning-Based FPGA Function Block Detection Method with Bitstream to Image Transformation
- Regula Sub-rosa: Latent Backdoor Attacks on Deep Neural Networks
- Fast acoustic scattering using convolutional neural networks
- A New Window Loss Function for Bone Fracture Detection and Localization in X-ray Images with Point-based Annotation
- Fast Object Detection in Compressed Video
- Multiple Instance Segmentation in Brachial Plexus Ultrasound Image Using BPMSegNet
- Team Delft's Robot Winner of the Amazon Picking Challenge 2016
- Data-Free Learning of Student Networks
- Reflective Decoding Network for Image Captioning
- Image Classification for Arabic: Assessing the Accuracy of Direct English to Arabic Translations
- Adaptive Reconstruction Network for Weakly Supervised Referring Expression Grounding
- Cops-Ref: A new Dataset and Task on Compositional Referring Expression Comprehension
- Revisiting Knowledge Distillation for Object Detection
- An Effective and Efficient Method for Detecting Hands in Egocentric Videos for Rehabilitation Applications
- Dual Recurrent Attention Units for Visual Question Answering
- Object-Aware Instance Labeling for Weakly Supervised Object Detection
- Deep execution monitor for robot assistive tasks
- 3D Context Enhanced Region-based Convolutional Neural Network for End-to-End Lesion Detection
- Generative Models as a Data Source for Multiview Representation Learning
- Artificial Immune System of Secure Face Recognition Against Adversarial Attacks
- Adaptive Object Detection with Dual Multi-Label Prediction
- Learning to Discretely Compose Reasoning Module Networks for Video Captioning
- Measuring Generalisation to Unseen Viewpoints, Articulations, Shapes and Objects for 3D Hand Pose Estimation under Hand-Object Interaction
- Self-Supervised Learning by Estimating Twin Class Distributions
- Boosting Weakly Supervised Object Detection with Progressive Knowledge Transfer
- No Peeking through My Windows: Conserving Privacy in Personal Drones
- Training Compact CNNs for Image Classification using Dynamic-coded Filter Fusion
- ScaleNAS: One-Shot Learning of Scale-Aware Representations for Visual Recognition
- Weakly-Supervised Object Detection Learning through Human-Robot Interaction
- Hire-MLP: Vision MLP via Hierarchical Rearrangement
- End-to-end detection-segmentation network with ROI convolution
- Generating Question Relevant Captions to Aid Visual Question Answering
- Diverse Sample Generation: Pushing the Limit of Generative Data-free Quantization
- Image Captioning based on Deep Learning Methods: A Survey
- LightningDOT: Pre-training Visual-Semantic Embeddings for Real-Time Image-Text Retrieval
- Robust Real-time Pedestrian Detection in Aerial Imagery on Jetson TX2
- Interpreting Adversarial Examples with Attributes
- Faster ILOD: Incremental Learning for Object Detectors based on Faster RCNN
- 360-Indoor: Towards Learning Real-World Objects in 360° Indoor Equirectangular Images
- Temporal-Channel Transformer for 3D Lidar-Based Video Object Detection in Autonomous Driving
- Motion-Excited Sampler: Video Adversarial Attack with Sparked Prior
- Context-Aware Single-Shot Detector
- Simultaneous Region Localization and Hash Coding for Fine-grained Image Retrieval
- FC2RN: A Fully Convolutional Corner Refinement Network for Accurate Multi-Oriented Scene Text Detection
- Synthetic Data Generation and Adaption for Object Detection in Smart Vending Machines
- Shape2Motion: Joint Analysis of Motion Parts and Attributes from 3D Shapes
- Three Branches: Detecting Actions With Richer Features
- THIA: Accelerating Video Analytics using Early Inference and Fine-Grained Query Planning
- Automatic Detection of Cardiac Chambers Using an Attention-based YOLOv4 Framework from Four-chamber View of Fetal Echocardiography
- The 1st Tiny Object Detection Challenge:Methods and Results
- Towards Unsupervised Crowd Counting via Regression-Detection Bi-knowledge Transfer
- Content-based Analysis of the Cultural Differences between TikTok and Douyin
- Cooling-Shrinking Attack: Blinding the Tracker with Imperceptible Noises
- Multi-Modal Attention-based Fusion Model for Semantic Segmentation of RGB-Depth Images
- Grounding Object Detections With Transcriptions
- STOA-VLP: Spatial-Temporal Modeling of Object and Action for Video-Language Pre-training
- Deep Learning in Computer-Aided Diagnosis and Treatment of Tumors: A Survey
- Loss re-scaling VQA: Revisiting the LanguagePrior Problem from a Class-imbalance View
- POMP: Pomcp-based Online Motion Planning for active visual search in indoor environments
- Latent Compositional Representations Improve Systematic Generalization in Grounded Question Answering
- End-to-End Video Object Detection with Spatial-Temporal Transformers
- Road User Detection in Videos
- Simple online and real-time tracking with occlusion handling
- FIBA: Frequency-Injection based Backdoor Attack in Medical Image Analysis
- Object Detection in Videos by High Quality Object Linking
- Robust Object Detection under Occlusion with Context-Aware CompositionalNets
- Exploring the Hierarchy in Relation Labels for Scene Graph Generation
- Amodal Segmentation Based on Visible Region Segmentation and Shape Prior
- Elephants Don't Pack Groceries: Robot Task Planning for Low Entropy Belief States
- NeuNetS: An Automated Synthesis Engine for Neural Network Design
- Domain-Specific Priors and Meta Learning for Few-Shot First-Person Action Recognition
- Improving Visual Question Answering by Referring to Generated Paragraph Captions
- Streaming Object Detection for 3-D Point Clouds
- SynthText3D: Synthesizing Scene Text Images from 3D Virtual Worlds
- Learning 3D Shapes as Multi-Layered Height-maps using 2D Convolutional Networks
- Can Synthetic Data Improve Object Detection Results for Remote Sensing Images?
- Zoom Out-and-In Network with Map Attention Decision for Region Proposal and Object Detection
- Learning to Generate Synthetic Data via Compositing
- Co-mining: Self-Supervised Learning for Sparsely Annotated Object Detection
- Context and Attribute Grounded Dense Captioning
- Segmenting Unknown 3D Objects from Real Depth Images using Mask R-CNN Trained on Synthetic Data
- On the Privacy Risks of Cell-Based NAS Architectures
- SoDA: Multi-Object Tracking with Soft Data Association
- A Context-and-Spatial Aware Network for Multi-Person Pose Estimation
- Localization Distillation for Dense Object Detection
- Deep Structured Feature Networks for Table Detection and Tabular Data Extraction from Scanned Financial Document Images
- A Multi-task Convolutional Neural Network for Autonomous Robotic Grasping in Object Stacking Scenes
- Continual Representation Learning for Biometric Identification
- Atherosclerotic carotid plaques on panoramic imaging: an automatic detection using deep learning with small dataset
- Jo-SRC: A Contrastive Approach for Combating Noisy Labels
- Human Extraction and Scene Transition utilizing Mask R-CNN
- CellLineNet: End-to-End Learning and Transfer Learning For Multiclass Epithelial Breast cell Line Classification via a Convolutional Neural Network
- Spatio-Temporal Action Detection with Multi-Object Interaction
- Learning Student Networks via Feature Embedding
- Small Towers Make Big Differences
- Lightweight Real-time Makeup Try-on in Mobile Browsers with Tiny CNN Models for Facial Tracking
- CAMERAS: Enhanced Resolution And Sanity preserving Class Activation Mapping for image saliency
- MetaAlign: Coordinating Domain Alignment and Classification for Unsupervised Domain Adaptation
- Answer Them All! Toward Universal Visual Question Answering Models
- A Multi-Stage Attentive Transfer Learning Framework for Improving COVID-19 Diagnosis
- Exploring Multi-Branch and High-Level Semantic Networks for Improving Pedestrian Detection
- COVID-19 personal protective equipment detection using real-time deep learning methods
- Optical Music Recognition: State of the Art and Major Challenges
- Epipolar-Guided Deep Object Matching for Scene Change Detection
- Asynchronous Interaction Aggregation for Action Detection
- Towards Precise End-to-end Weakly Supervised Object Detection Network
- Bilinear Graph Networks for Visual Question Answering
- Towards General Purpose Vision Systems
- 2nd Place Solution for Waymo Open Dataset Challenge -- 2D Object Detection
- Context-Aware Crowd Counting
- Unpaired Pose Guided Human Image Generation
- Face Parsing with RoI Tanh-Warping
- Residual Attention based Network for Hand Bone Age Assessment
- SWA Object Detection
- Measuring economic activity from space: a case study using flying airplanes and COVID-19
- Neural Architecture Search by Estimation of Network Structure Distributions
- Object Recognition with and without Objects
- Towards Generalized and Incremental Few-Shot Object Detection
- Mask2CAD: 3D Shape Prediction by Learning to Segment and Retrieve
- Explainable Goal-Driven Agents and Robots -- A Comprehensive Review
- Learning to Exploit Multiple Vision Modalities by Using Grafted Networks
- Towards High-Quality Temporal Action Detection with Sparse Proposals
- EMPNet: Neural Localisation and Mapping Using Embedded Memory Points
- PAFNet: An Efficient Anchor-Free Object Detector Guidance
- Condensation-Net: Memory-Efficient Network Architecture with Cross-Channel Pooling Layers and Virtual Feature Maps
- Chinese Street View Text: Large-scale Chinese Text Reading with Partially Supervised Learning
- General Instance Distillation for Object Detection
- RILOD: Near Real-Time Incremental Learning for Object Detection at the Edge
- Adversarially Trained Object Detector for Unsupervised Domain Adaptation
- Towards Multi-class Object Detection in Unconstrained Remote Sensing Imagery
- Straight to Shapes: Real-time Detection of Encoded Shapes
- Deep Learning Based Instance Segmentation in 3D Biomedical Images Using Weak Annotation
- TextCohesion: Detecting Text for Arbitrary Shapes
- One-Shot Unsupervised Cross-Domain Detection
- A novel integrated industrial approach with cobots in the age of industry 4.0 through conversational interaction and computer vision
- Generative Zero-shot Network Quantization
- Lidar Panoptic Segmentation in an Open World
- Globally-Aware Multiple Instance Classifier for Breast Cancer Screening
- Deep Instance Segmentation and Visual Servoing to Play Jenga with a Cost-Effective Robotic System
- P2B: Point-to-Box Network for 3D Object Tracking in Point Clouds
- A Unified Hardware Architecture for Convolutions and Deconvolutions in CNN
- Semantic Feature Matching for Robust Mapping in Agriculture
- Resilience of Autonomous Vehicle Object Category Detection to Universal Adversarial Perturbations
- Deep Similarity Metric Learning for Real-Time Pedestrian Tracking
- DeRPN: Taking a further step toward more general object detection
- Unsupervised Deep Feature Transfer for Low Resolution Image Classification
- DeepACC:Automate Chromosome Classification based on Metaphase Images using Deep Learning Framework Fused with Prior Knowledge
- Conv-MPN: Convolutional Message Passing Neural Network for Structured Outdoor Architecture Reconstruction
- Monitoring spatial sustainable development: Semi-automated analysis of satellite and aerial images for energy transition and sustainability indicators
- Fire SSD: Wide Fire Modules based Single Shot Detector on Edge Device
- Retrieve, Read, Rerank: Towards End-to-End Multi-Document Reading Comprehension
- Structural Pruning in Deep Neural Networks: A Small-World Approach
- Detecting retail products in situ using CNN without human effort labeling
- iffDetector: Inference-aware Feature Filtering for Object Detection
- Delving into Robust Object Detection from Unmanned Aerial Vehicles: A Deep Nuisance Disentanglement Approach
- On the Robustness of Human Pose Estimation
- Diagnosing Error in Temporal Action Detectors
- Artificial Intelligence For Breast Cancer Detection: Trends & Directions
- RefineMask: Towards High-Quality Instance Segmentation with Fine-Grained Features
- GestARLite: An On-Device Pointing Finger Based Gestural Interface for Smartphones and Video See-Through Head-Mounts
- Ego-Exo: Transferring Visual Representations from Third-person to First-person Videos
- ACP++: Action Co-occurrence Priors for Human-Object Interaction Detection
- Image Resizing by Reconstruction from Deep Features
- KIT MOMA: A Mobile Machines Dataset
- Adversarial Parameter Defense by Multi-Step Risk Minimization
- A Large Scale Urban Surveillance Video Dataset for Multiple-Object Tracking and Behavior Analysis
- IAN: Combining Generative Adversarial Networks for Imaginative Face Generation
- Greedy Gradient Ensemble for Robust Visual Question Answering
- V2F-Net: Explicit Decomposition of Occluded Pedestrian Detection
- RigNet: Repetitive Image Guided Network for Depth Completion
- Progressive Cluster Purification for Unsupervised Feature Learning
- A Novel Graph-based Multi-modal Fusion Encoder for Neural Machine Translation
- Channel Locality Block: A Variant of Squeeze-and-Excitation
- Curved Text Detection in Natural Scene Images with Semi- and Weakly-Supervised Learning
- Domain-Specific Suppression for Adaptive Object Detection
- Split and Connect: A Universal Tracklet Booster for Multi-Object Tracking
- The Mapillary Traffic Sign Dataset for Detection and Classification on a Global Scale
- Joint Face Detection and Facial Motion Retargeting for Multiple Faces
- Segmentation and Recovery of Superquadric Models using Convolutional Neural Networks
- Journalistic Guidelines Aware News Image Captioning
- All You Need Is Boundary: Toward Arbitrary-Shaped Text Spotting
- Cascaded Subpatch Networks for Effective CNNs
- I3Net: Implicit Instance-Invariant Network for Adapting One-Stage Object Detectors
- Pseudo-IoU: Improving Label Assignment in Anchor-Free Object Detection
- Connecting the Dots: Detecting Adversarial Perturbations Using Context Inconsistency
- SGNet: A Super-class Guided Network for Image Classification and Object Detection
- Optimizing Through Learned Errors for Accurate Sports Field Registration
- OpenViDial 2.0: A Larger-Scale, Open-Domain Dialogue Generation Dataset with Visual Contexts
- Democratizing Production-Scale Distributed Deep Learning
- IoU Attack: Towards Temporally Coherent Black-Box Adversarial Attack for Visual Object Tracking
- IIIT-AR-13K: A New Dataset for Graphical Object Detection in Documents
- Looking Beyond Two Frames: End-to-End Multi-Object Tracking Using Spatial and Temporal Transformers
- Modality-Balanced Models for Visual Dialogue
- Exploiting Playbacks in Unsupervised Domain Adaptation for 3D Object Detection
- GRIP: Generative Robust Inference and Perception for Semantic Robot Manipulation in Adversarial Environments
- Shift Equivariance in Object Detection
- Uncertainty-Aware Unsupervised Domain Adaptation in Object Detection
- Learning Gating ConvNet for Two-Stream based Methods in Action Recognition
- 1st Place Solution for the UVO Challenge on Image-based Open-World Segmentation 2021
- Weakly supervised cross-domain alignment with optimal transport
- Robustness in Compressed Neural Networks for Object Detection
- Design and Interpretation of Universal Adversarial Patches in Face Detection
- CCL: Cross-modal Correlation Learning with Multi-grained Fusion by Hierarchical Network
- Geometry-Aware Video Object Detection for Static Cameras
- A Possible Reason for why Data-Driven Beats Theory-Driven Computer Vision
- Re-labeling ImageNet: from Single to Multi-Labels, from Global to Localized Labels
- Probabilistic Embeddings for Cross-Modal Retrieval
- Transferring Domain-Agnostic Knowledge in Video Question Answering
- Improving Deep Lesion Detection Using 3D Contextual and Spatial Attention
- Structure Aware SLAM using Quadrics and Planes
- Beyond Point Clouds: A Knowledge-Aided High Resolution Imaging Radar Deep Detector for Autonomous Driving
- OPANAS: One-Shot Path Aggregation Network Architecture Search for Object Detection
- Towards Calibrated Model for Long-Tailed Visual Recognition from Prior Perspective
- CDeC-Net: Composite Deformable Cascade Network for Table Detection in Document Images
- Fast Region Proposal Learning for Object Detection for Robotics
- Grab: Fast and Accurate Sensor Processing for Cashier-Free Shopping
- Class-specific Anchoring Proposal for 3D Object Recognition in LIDAR and RGB Images
- Semantic Segmentation from Limited Training Data
- Scene-Graph Augmented Data-Driven Risk Assessment of Autonomous Vehicle Decisions
- Teaching Robots Novel Objects by Pointing at Them
- Guiding the Creation of Deep Learning-based Object Detectors
- DeepApple: Deep Learning-based Apple Detection using a Suppression Mask R-CNN
- Dynamic Edge Weights in Graph Neural Networks for 3D Object Detection
- Backbone Can Not be Trained at Once: Rolling Back to Pre-trained Network for Person Re-Identification
- Collaborative Training between Region Proposal Localization and Classification for Domain Adaptive Object Detection
- Towards Overcoming False Positives in Visual Relationship Detection
- Mixed-Precision Quantized Neural Network with Progressively Decreasing Bitwidth For Image Classification and Object Detection
- Spatial-Temporal Block and LSTM Network for Pedestrian Trajectories Prediction
- Weakly Supervised Object Boundaries
- Semi-Supervised Image Deraining using Gaussian Processes
- Shape-Texture Debiased Neural Network Training
- Deep Regionlets: Blended Representation and Deep Learning for Generic Object Detection
- MixSearch: Searching for Domain Generalized Medical Image Segmentation Architectures
- Value of Temporal Dynamics Information in Driving Scene Segmentation
- Bi-Classifier Determinacy Maximization for Unsupervised Domain Adaptation
- Automated Classification of Helium Ingress in Irradiated X-750
- Regularizing Attention Networks for Anomaly Detection in Visual Question Answering
- Constrained R-CNN: A general image manipulation detection model
- Relation Transformer Network
- Information Bottleneck Constrained Latent Bidirectional Embedding for Zero-Shot Learning
- Self-supervised pre-training and contrastive representation learning for multiple-choice video QA
- Attribute-aware Pedestrian Detection in a Crowd
- Seesaw Loss for Long-Tailed Instance Segmentation
- Mixed Supervised Object Detection with Robust Objectness Transfer
- American Sign Language fingerspelling recognition in the wild
- Depthwise Spatio-Temporal STFT Convolutional Neural Networks for Human Action Recognition
- HWNet v2: An Efficient Word Image Representation for Handwritten Documents
- Deep Multi-camera People Detection
- A Solution to Product detection in Densely Packed Scenes
- Pixel and Feature Level Based Domain Adaption for Object Detection in Autonomous Driving
- Visual Grounding Methods for VQA are Working for the Wrong Reasons!
- Unsupervised Pre-training for Person Re-identification
- DiscoBox: Weakly Supervised Instance Segmentation and Semantic Correspondence from Box Supervision
- Learning Gaussian Maps for Dense Object Detection
- SiamMOT: Siamese Multi-Object Tracking
- WSSOD: A New Pipeline for Weakly- and Semi-Supervised Object Detection
- Multimodal Learning for Hateful Memes Detection
- PSC-Net: Learning Part Spatial Co-occurrence for Occluded Pedestrian Detection
- Dont Even Look Once: Synthesizing Features for Zero-Shot Detection
- Detect-and-describe: Joint learning framework for detection and description of objects
- Unsupervised Multimodal Neural Machine Translation with Pseudo Visual Pivoting
- Privid: Practical, Privacy-Preserving Video Analytics Queries
- Manipulating Identical Filter Redundancy for Efficient Pruning on Deep and Complicated CNN
- Re-ID Driven Localization Refinement for Person Search
- The Importance and the Limitations of Sim2Real for Robotic Manipulation in Precision Agriculture
- Enhancing a Neurocognitive Shared Visuomotor Model for Object Identification, Localization, and Grasping With Learning From Auxiliary Tasks
- QEBA: Query-Efficient Boundary-Based Blackbox Attack
- iReason: Multimodal Commonsense Reasoning using Videos and Natural Language with Interpretability
- An FPGA-Accelerated Design for Deep Learning Pedestrian Detection in Self-Driving Vehicles
- Exploiting the Inherent Limitation of L0 Adversarial Examples
- The Devil is in Classification: A Simple Framework for Long-tail Object Detection and Instance Segmentation
- Seeing is Knowing! Fact-based Visual Question Answering using Knowledge Graph Embeddings
- A multi-task convolutional neural network for mega-city analysis using very high resolution satellite imagery and geospatial data
- 3DV: 3D Dynamic Voxel for Action Recognition in Depth Video
- Deep Miner: A Deep and Multi-branch Network which Mines Rich and Diverse Features for Person Re-identification
- Deep ChArUco: Dark ChArUco Marker Pose Estimation
- Self-Supervised Representation Learning for Visual Anomaly Detection
- Unsupervised Instance Segmentation in Microscopy Images via Panoptic Domain Adaptation and Task Re-weighting
- Symmetry and Group in Attribute-Object Compositions
- A Picture May Be Worth a Hundred Words for Visual Question Answering
- Convolutional Oriented Boundaries: From Image Segmentation to High-Level Tasks
- Concatenated Feature Pyramid Network for Instance Segmentation
- Density Map Guided Object Detection in Aerial Images
- Multi-Loss Sub-Ensembles for Accurate Classification with Uncertainty Estimation
- A Structured Model For Action Detection
- Activity Driven Weakly Supervised Object Detection
- Instance-Level Relative Saliency Ranking with Graph Reasoning
- Detect-to-Retrieve: Efficient Regional Aggregation for Image Search
- Incremental Learning for Robot Perception through HRI
- Domain Adaptation and Image Classification via Deep Conditional Adaptation Network
- MatrixNets: A New Scale and Aspect Ratio Aware Architecture for Object Detection
- MADAN: Multi-source Adversarial Domain Aggregation Network for Domain Adaptation
- Learning to Fuse Asymmetric Feature Maps in Siamese Trackers
- ReLaText: Exploiting Visual Relationships for Arbitrary-Shaped Scene Text Detection with Graph Convolutional Networks
- On Robustness of Lane Detection Models to Physical-World Adversarial Attacks in Autonomous Driving
- Detecting Small, Densely Distributed Objects with Filter-Amplifier Networks and Loss Boosting
- 3D Feature Pyramid Attention Module for Robust Visual Speech Recognition
- Heterogeneous Contrastive Learning: Encoding Spatial Information for Compact Visual Representations
- Faceness-Net: Face Detection through Deep Facial Part Responses
- Gaining Extra Supervision via Multi-task learning for Multi-Modal Video Question Answering
- Learning About Objects by Learning to Interact with Them
- 1st place solution for AVA-Kinetics Crossover in AcitivityNet Challenge 2020
- Discriminative Appearance Modeling with Multi-track Pooling for Real-time Multi-object Tracking
- A Large-Scale Dataset for Benchmarking Elevator Button Segmentation and Character Recognition
- Augmented Parallel-Pyramid Net for Attention Guided Pose-Estimation
- Cheaper Pre-training Lunch: An Efficient Paradigm for Object Detection
- Learning Effective Visual Relationship Detector on 1 GPU
- The iMaterialist Fashion Attribute Dataset
- Applications of Deep Learning in Fundus Images: A Review
- Active Perception for Ambiguous Objects Classification
- GTNet:Guided Transformer Network for Detecting Human-Object Interactions
- PnP-DETR: Towards Efficient Visual Analysis with Transformers
- Dataflow-based Joint Quantization of Weights and Activations for Deep Neural Networks
- Semantic Compositional Learning for Low-shot Scene Graph Generation
- ML-EXray: Visibility into ML Deployment on the Edge
- Real-time Embedded Person Detection and Tracking for Shopping Behaviour Analysis
- No Surprises: Training Robust Lung Nodule Detection for Low-Dose CT Scans by Augmenting with Adversarial Attacks
- Decoupling Localization and Classification in Single Shot Temporal Action Detection
- Text Perceptron: Towards End-to-End Arbitrary-Shaped Text Spotting
- Review: deep learning on 3D point clouds
- Ship Instance Segmentation From Remote Sensing Images Using Sequence Local Context Module
- SaccadeNet: A Fast and Accurate Object Detector
- Dynamically Throttleable Neural Networks (TNN)
- m-RevNet: Deep Reversible Neural Networks with Momentum
- Joint Layout Analysis, Character Detection and Recognition for Historical Document Digitization
- Unsupervised Domain Adaptation for Multispectral Pedestrian Detection
- Moment-Based Domain Adaptation: Learning Bounds and Algorithms
- Learning from Web Data: the Benefit of Unsupervised Object Localization
- Whole-Body Human Pose Estimation in the Wild
- EventNet: Asynchronous Recursive Event Processing
- Unsupervised Temporal Feature Aggregation for Event Detection in Unstructured Sports Videos
- Channel-wise Alignment for Adaptive Object Detection
- Multi-View Multi-Person 3D Pose Estimation with Plane Sweep Stereo
- Cloud-based Image Classification Service Is Not Robust To Simple Transformations: A Forgotten Battlefield
- Rethinking Drone-Based Search and Rescue with Aerial Person Detection
- REGNet: REgion-based Grasp Network for End-to-end Grasp Detection in Point Clouds
- Mixture of Pre-processing Experts Model for Noise Robust Deep Learning on Resource Constrained Platforms
- Pointly-Supervised Instance Segmentation
- BUZz: BUffer Zones for defending adversarial examples in image classification
- VIN: Voxel-based Implicit Network for Joint 3D Object Detection and Segmentation for Lidars
- PerMO: Perceiving More at Once from a Single Image for Autonomous Driving
- Skin disease diagnosis with deep learning: a review
- Learning Image Aesthetic Assessment from Object-level Visual Components
- RAPiD: Rotation-Aware People Detection in Overhead Fisheye Images
- Curriculum By Smoothing
- Hierarchy Parsing for Image Captioning
- VisualMRC: Machine Reading Comprehension on Document Images
- Deep Learning for Automatic Quality Grading of Mangoes: Methods and Insights
- Deeply Aligned Adaptation for Cross-domain Object Detection
- A Parameterized Approach to Personalized Variable Length Summarization of Soccer Matches
- GLSD: The Global Large-Scale Ship Database and Baseline Evaluations
- Segmentation Mask Guided End-to-End Person Search
- 3D Aggregated Faster R-CNN for General Lesion Detection
- Scale Invariant Fully Convolutional Network: Detecting Hands Efficiently
- Learning to Filter: Siamese Relation Network for Robust Tracking
- Attention-Driven Dynamic Graph Convolutional Network for Multi-Label Image Recognition
- Improving Contrastive Learning by Visualizing Feature Transformation
- FNA++: Fast Network Adaptation via Parameter Remapping and Architecture Search
- Multi-Class 3D Object Detection Within Volumetric 3D Computed Tomography Baggage Security Screening Imagery
- 3D-DETNet: a Single Stage Video-Based Vehicle Detector
- Edited Media Understanding Frames: Reasoning About the Intent and Implications of Visual Misinformation
- Modeling Spatio-Temporal Human Track Structure for Action Localization
- Farm land weed detection with region-based deep convolutional neural networks
- Long-distance tiny face detection based on enhanced YOLOv3 for unmanned system
- Fingerspelling recognition in the wild with iterative visual attention
- Let There be Light: Improved Traffic Surveillance via Detail Preserving Night-to-Day Transfer
- Deep Blur Mapping: Exploiting High-Level Semantics by Deep Neural Networks
- POI: Multiple Object Tracking with High Performance Detection and Appearance Feature
- DoReMi: First glance at a universal OMR dataset
- SS3D: Single Shot 3D Object Detector
- 3D Pose Estimation for Fine-Grained Object Categories
- Data-Driven Vehicle Trajectory Forecasting
- Visual Attention Model for Cross-sectional Stock Return Prediction and End-to-End Multimodal Market Representation Learning
- SAIA: Split Artificial Intelligence Architecture for Mobile Healthcare System
- A Review on Object Pose Recovery: from 3D Bounding Box Detectors to Full 6D Pose Estimators
- Human Following for Wheeled Robot with Monocular Pan-tilt Camera
- Geometry-Aware Instance Segmentation with Disparity Maps
- 3D Crowd Counting via Geometric Attention-guided Multi-View Fusion
- Fashionpedia: Ontology, Segmentation, and an Attribute Localization Dataset
- Scale Matters: Temporal Scale Aggregation Network for Precise Action Localization in Untrimmed Videos
- MixModule: Mixed CNN Kernel Module for Medical Image Segmentation
- Multi-scale Aggregation R-CNN for 2D Multi-person Pose Estimation
- Dense Relational Image Captioning via Multi-task Triple-Stream Networks
- Mask is All You Need: Rethinking Mask R-CNN for Dense and Arbitrary-Shaped Scene Text Detection
- Mixup Regularization for Region Proposal based Object Detectors
- Context-Aware RCNN: A Baseline for Action Detection in Videos
- Visually Grounded Continual Learning of Compositional Phrases
- Multi-Scale 2D Temporal Adjacent Networks for Moment Localization with Natural Language
- Straight to Shapes++: Real-time Instance Segmentation Made More Accurate
- Non-Local Context Encoder: Robust Biomedical Image Segmentation against Adversarial Attacks
- Efficient Global Multi-object Tracking Under Minimum-cost Circulation Framework
- Over-the-Air Adversarial Flickering Attacks against Video Recognition Networks
- Layerwise Optimization by Gradient Decomposition for Continual Learning
- Learning to Separate: Detecting Heavily-Occluded Objects in Urban Scenes
- Robust 6D Object Pose Estimation by Learning RGB-D Features
- Label-Attention Transformer with Geometrically Coherent Objects for Image Captioning
- Neural Architecture Refinement: A Practical Way for Avoiding Overfitting in NAS
- Detection as Regression: Certified Object Detection by Median Smoothing
- Stereo Vision Based Single-Shot 6D Object Pose Estimation for Bin-Picking by a Robot Manipulator
- Integrating Image Captioning with Rule-based Entity Masking
- Exploring the Capacity of an Orderless Box Discretization Network for Multi-orientation Scene Text Detection
- Distribution-sensitive Information Retention for Accurate Binary Neural Network
- Point Proposal Network: Accelerating Point Source Detection Through Deep Learning
- Crossover Learning for Fast Online Video Instance Segmentation
- Low-Latency Human Action Recognition with Weighted Multi-Region Convolutional Neural Network
- Multi-label Detection and Classification of Red Blood Cells in Microscopic Images
- StyleAugment: Learning Texture De-biased Representations by Style Augmentation without Pre-defined Textures
- SAVERS: SAR ATR with Verification Support Based on Convolutional Neural Network
- Learning a Domain Classifier Bank for Unsupervised Adaptive Object Detection
- Towards Generalization and Data Efficient Learning of Deep Robotic Grasping
- Representing Videos as Discriminative Sub-graphs for Action Recognition
- Joint Representation and Truncated Inference Learning for Correlation Filter based Tracking
- Exemplar-Based Open-Set Panoptic Segmentation Network
- Noisy Annotation Refinement for Object Detection
- Rethinking the Artificial Neural Networks: A Mesh of Subnets with a Central Mechanism for Storing and Predicting the Data
- Cyclic Co-Learning of Sounding Object Visual Grounding and Sound Separation
- IQDet: Instance-wise Quality Distribution Sampling for Object Detection
- AIBench Training: Balanced Industry-Standard AI Training Benchmarking
- Robustness Enhancement of Object Detection in Advanced Driver Assistance Systems (ADAS)
- Edge-Cloud Collaborated Object Detection via Difficult-Case Discriminator
- DeepSEED: 3D Squeeze-and-Excitation Encoder-Decoder Convolutional Neural Networks for Pulmonary Nodule Detection
- SASL: Saliency-Adaptive Sparsity Learning for Neural Network Acceleration
- Mi YouTube es Su YouTube? Analyzing the Cultures using YouTube Thumbnails of Popular Videos
- Social Adaptive Module for Weakly-supervised Group Activity Recognition
- High-order Tensor Pooling with Attention for Action Recognition
- Adaptive Offline Quintuplet Loss for Image-Text Matching
- Learning to Count Objects with Few Exemplar Annotations
- FrostNet: Towards Quantization-Aware Network Architecture Search
- CRACT: Cascaded Regression-Align-Classification for Robust Visual Tracking
- DAWN: Dual Augmented Memory Network for Unsupervised Video Object Tracking
- Two-Stream Region Convolutional 3D Network for Temporal Activity Detection
- A^2-FPN: Attention Aggregation based Feature Pyramid Network for Instance Segmentation
- Tell Me What They're Holding: Weakly-supervised Object Detection with Transferable Knowledge from Human-object Interaction
- AdaVQA: Overcoming Language Priors with Adapted Margin Cosine Loss
- Class-Aware Robust Adversarial Training for Object Detection
- Membership Inference Attacks Against Object Detection Models
- Jigsaw Clustering for Unsupervised Visual Representation Learning
- Deformable Tube Network for Action Detection in Videos
- Optimal Gradient Checkpoint Search for Arbitrary Computation Graphs
- Soybean pod and seed counting in both outdoor fields and indoor laboratories using unions of deep neural networks
- Investigating Attention Mechanism in 3D Point Cloud Object Detection
- Translate-to-Recognize Networks for RGB-D Scene Recognition
- A Robotic Approach towards Quantifying Epipelagic Bound Plastic Using Deep Visual Models
- Multiple Object Tracking with Motion and Appearance Cues
- Towards Automatic Construction of Diverse, High-quality Image Dataset
- Decoder Choice Network for Meta-Learning
- Overcoming Statistical Shortcuts for Open-ended Visual Counting
- Multi-Layer Content Interaction Through Quaternion Product For Visual Question Answering
- Soft Prototyping Camera Designs for Car Detection Based on a Convolutional Neural Network
- When We First Met: Visual-Inertial Person Localization for Co-Robot Rendezvous
- GTNet: Generative Transfer Network for Zero-Shot Object Detection
- Deep Multi-task Learning for Facial Expression Recognition and Synthesis Based on Selective Feature Sharing
- Self-supervised Robust Object Detectors from Partially Labelled Datasets
- ACP: Automatic Channel Pruning via Clustering and Swarm Intelligence Optimization for CNN
- Learning Channel Inter-dependencies at Multiple Scales on Dense Networks for Face Recognition
- Exploiting Contextual Information with Deep Neural Networks
- Discrete-continuous Action Space Policy Gradient-based Attention for Image-Text Matching
- Motorcycle detection and classification in urban Scenarios using a model based on Faster R-CNN
- CRIC: A VQA Dataset for Compositional Reasoning on Vision and Commonsense
- Developing a Compressed Object Detection Model based on YOLOv4 for Deployment on Embedded GPU Platform of Autonomous System
- Generic Tubelet Proposals for Action Localization
- Learning 3D-aware Egocentric Spatial-Temporal Interaction via Graph Convolutional Networks
- Chargrid-OCR: End-to-end Trainable Optical Character Recognition for Printed Documents using Instance Segmentation
- Image-based monitoring of bolt loosening through deep-learning-based integrated detection and tracking
- Captioning Images with Novel Objects via Online Vocabulary Expansion
- Joint Multimedia Event Extraction from Video and Article
- Quantum-soft QUBO Suppression for Accurate Object Detection
- Progress Regression RNN for Online Spatial-Temporal Action Localization in Unconstrained Videos
- Efficient Adversarial Attacks for Visual Object Tracking
- A Selective Survey on Versatile Knowledge Distillation Paradigm for Neural Network Models
- Reconstruction of Simulation-Based Physical Field by Reconstruction Neural Network Method
- Data Priming Network for Automatic Check-Out
- ORBIT: A Real-World Few-Shot Dataset for Teachable Object Recognition
- Fusing Saliency Maps with Region Proposals for Unsupervised Object Localization
- Pretraining Techniques for Sequence-to-Sequence Voice Conversion
- iShape: A First Step Towards Irregular Shape Instance Segmentation
- Reviewing continual learning from the perspective of human-level intelligence
- SID: Incremental Learning for Anchor-Free Object Detection via Selective and Inter-Related Distillation
- User Constrained Thumbnail Generation using Adaptive Convolutions
- Context-Aware Unsupervised Clustering for Person Search
- A Unified Object Motion and Affinity Model for Online Multi-Object Tracking
- Relationship-Embedded Representation Learning for Grounding Referring Expressions
- TextRay: Contour-based Geometric Modeling for Arbitrary-shaped Scene Text Detection
- FarSee-Net: Real-Time Semantic Segmentation by Efficient Multi-scale Context Aggregation and Feature Space Super-resolution
- AACP: Model Compression by Accurate and Automatic Channel Pruning
- Training Generative Adversarial Networks in One Stage
- Learning to Discriminate Information for Online Action Detection
- CADP: A Novel Dataset for CCTV Traffic Camera based Accident Analysis
- Towards Panoptic 3D Parsing for Single Image in the Wild
- Accurate Anchor Free Tracking
- A Multi-task Contextual Atrous Residual Network for Brain Tumor Detection & Segmentation
- HALP: Hardware-Aware Latency Pruning
- Skeleton Based Action Recognition using a Stacked Denoising Autoencoder with Constraints of Privileged Information
- Beyond the Camera: Neural Networks in World Coordinates
- Object Detection with a Unified Label Space from Multiple Datasets
- A Frank-Wolfe Framework for Efficient and Effective Adversarial Attacks
- Real-Time Visual Object Tracking via Few-Shot Learning
- Towards a Framework for Visual Intelligence in Service Robotics: Epistemic Requirements and Gap Analysis
- EBoWs: An End-to-End Bag-of-Words Model via Deep Convolutional Neural Network
- StampNet: unsupervised multi-class object discovery
- Show, Control and Tell: A Framework for Generating Controllable and Grounded Captions
- OGNet: Salient Object Detection with Output-guided Attention Module
- Multilevel Language and Vision Integration for Text-to-Clip Retrieval
- Understanding the effects of artifacts on automated polyp detection and incorporating that knowledge via learning without forgetting
- Grasping Detection Network with Uncertainty Estimation for Confidence-Driven Semi-Supervised Domain Adaptation
- ASAP-NMS: Accelerating Non-Maximum Suppression Using Spatially Aware Priors
- X-LineNet: Detecting Aircraft in Remote Sensing Images by a pair of Intersecting Line Segments
- Scaling Object Detection by Transferring Classification Weights
- Context-Aware Zero-Shot Recognition
- Relationship Oriented Affordance Learning through Manipulation Graph Construction
- ROAM: Recurrently Optimizing Tracking Model
- Supervised Visual Attention for Simultaneous Multimodal Machine Translation
- Fast-Tracker 2.0: Improving Autonomy of Aerial Tracking with Active Vision and Human Location Regression
- Dance with Flow: Two-in-One Stream Action Detection
- Learning Sparse Mixture of Experts for Visual Question Answering
- Online Multi-Object Tracking with Dual Matching Attention Networks
- Fast Object Segmentation Learning with Kernel-based Methods for Robotics
- Integrating Objects into Monocular SLAM: Line Based Category Specific Models
- Pavement Distress Detection and Segmentation using YOLOv4 and DeepLabv3 on Pavements in the Philippines
- Visual Tracking via Dynamic Memory Networks
- Adaptive Control of Embedding Strength in Image Watermarking using Neural Networks
- Privacy-preserving Object Detection
- A Deep Multi-task Learning Approach to Skin Lesion Classification
- Towards Adversarially Robust Object Detection
- Real-time Visual Object Tracking with Natural Language Description
- SynthRef: Generation of Synthetic Referring Expressions for Object Segmentation
- Exploring Semantic Relationships for Unpaired Image Captioning
- AIBench Scenario: Scenario-distilling AI Benchmarking
- Perception Improvement for Free: Exploring Imperceptible Black-box Adversarial Attacks on Image Classification
- Adversarial Augmentation for Enhancing Classification of Mammography Images
- Object-aware Feature Aggregation for Video Object Detection
- LAMP: Label Augmented Multimodal Pretraining
- Reverse-engineering Bar Charts Using Neural Networks
- F-Siamese Tracker: A Frustum-based Double Siamese Network for 3D Single Object Tracking
- A Mask R-CNN approach to counting bacterial colony forming units in pharmaceutical development
- Quantifying and Alleviating the Language Prior Problem in Visual Question Answering
- Gated Multi-layer Convolutional Feature Extraction Network for Robust Pedestrian Detection
- COCAS: A Large-Scale Clothes Changing Person Dataset for Re-identification
- Transferable Active Grasping and Real Embodied Dataset
- CAD-PU: A Curvature-Adaptive Deep Learning Solution for Point Set Upsampling
- Rethinking Open-Set Object Detection: Issues, a New Formulation, and Taxonomy
- Progressive Learning of Low-Precision Networks
- SECS: Efficient Deep Stream Processing via Class Skew Dichotomy
- Dynamic Filtering with Large Sampling Field for ConvNets
- Training and Testing Object Detectors with Virtual Images
- AIO-P: Expanding Neural Performance Predictors Beyond Image Classification
- Improving Face Detection Performance with 3D-Rendered Synthetic Data
- Gotta Adapt 'Em All: Joint Pixel and Feature-Level Domain Adaptation for Recognition in the Wild
- An End-to-End Foreground-Aware Network for Person Re-Identification
- PC2WF: 3D Wireframe Reconstruction from Raw Point Clouds
- Panoster: End-to-end Panoptic Segmentation of LiDAR Point Clouds
- SAFCAR: Structured Attention Fusion for Compositional Action Recognition
- Not All Words are Equal: Video-specific Information Loss for Video Captioning
- Joint-DetNAS: Upgrade Your Detector with NAS, Pruning and Dynamic Distillation
- Frustum VoxNet for 3D object detection from RGB-D or Depth images
- Multiple Anchor Learning for Visual Object Detection
- Object Detection for Understanding Assembly Instruction Using Context-aware Data Augmentation and Cascade Mask R-CNN
- Multi-Branch Fully Convolutional Network for Face Detection
- Deep learning for brake squeal: vibration detection, characterization and prediction
- Deeply Activated Salient Region for Instance Search
- Parameter Efficient Deep Neural Networks with Bilinear Projections
- A Taught-Obesrve-Ask (TOA) Method for Object Detection with Critical Supervision
- A study on using image based machine learning methods to develop the surrogate models of stamp forming simulations
- Instance-Level Task Parameters: A Robust Multi-task Weighting Framework
- EnTri: Ensemble Learning with Tri-level Representations for Explainable Scene Recognition
- WiderPerson: A Diverse Dataset for Dense Pedestrian Detection in the Wild
- Condition directed Multi-domain Adversarial Learning for Loop Closure Detection
- TextNet: Irregular Text Reading from Images with an End-to-End Trainable Network
- Act Like a Radiologist: Towards Reliable Multi-view Correspondence Reasoning for Mammogram Mass Detection
- A Simple yet Effective Baseline for Robust Deep Learning with Noisy Labels
- Propose-and-Attend Single Shot Detector
- Domain Adaptive SiamRPN++ for Object Tracking in the Wild
- Scale Optimization for Full-Image-CNN Vehicle Detection
- FRDet: Balanced and Lightweight Object Detector based on Fire-Residual Modules for Embedded Processor of Autonomous Driving
- Learning Orientation-Estimation Convolutional Neural Network for Building Detection in Optical Remote Sensing Image
- OneAdapt: Fast Configuration Adaptation for Video Analytics Applications via Backpropagation
- Autonomous Navigation in Dynamic Environments: Deep Learning-Based Approach
- Visual Semantic Information Pursuit: A Survey
- Robust Data Association for Object-level Semantic SLAM
- Towards Better Object Detection in Scale Variation with Adaptive Feature Selection
- Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck
- Phrase Localization Without Paired Training Examples
- GLiT: Neural Architecture Search for Global and Local Image Transformer
- Obj-GloVe: Scene-Based Contextual Object Embedding
- Efficient Pig Counting in Crowds with Keypoints Tracking and Spatial-aware Temporal Response Filtering
- CPM R-CNN: Calibrating Point-guided Misalignment in Object Detection
- Object Proposal with Kernelized Partial Ranking
- On the Compressive Power of Deep Rectifier Networks for High Resolution Representation of Class Boundaries
- Inner-Scene Similarities as a Contextual Cue for Object Detection
- Attentive CT Lesion Detection Using Deep Pyramid Inference with Multi-Scale Booster
- vireoJD-MM at Activity Detection in Extended Videos
- IMP: Instance Mask Projection for High Accuracy Semantic Segmentation of Things
- Learning Spatio-Temporal Representation with Local and Global Diffusion
- Spatio-temporal Video Re-localization by Warp LSTM
- High Frequency Residual Learning for Multi-Scale Image Classification
- A CNN-RNN Architecture for Multi-Label Weather Recognition
- Collaboration Analysis Using Deep Learning
- A Training-free, One-shot Detection Framework For Geospatial Objects In Remote Sensing Images
- Modularized Textual Grounding for Counterfactual Resilience
- LookUP: Vision-Only Real-Time Precise Underground Localisation for Autonomous Mining Vehicles
- Disentangled Deep Autoencoding Regularization for Robust Image Classification
- Background subtraction on depth videos with convolutional neural networks
- Automatic Surface Area and Volume Prediction on Ellipsoidal Ham using Deep Learning
- Detecting Text in the Wild with Deep Character Embedding Network
- URNet : User-Resizable Residual Networks with Conditional Gating Module
- Speaker Diarization with Region Proposal Network
- A Novel and Efficient Tumor Detection Framework for Pancreatic Cancer via CT Images
- Look, Read and Feel: Benchmarking Ads Understanding with Multimodal Multitask Learning
- Smart Home Appliances: Chat with Your Fridge
- TextTubes for Detecting Curved Text in the Wild
- Knowledge-Enriched Visual Storytelling
- Visual Dialogue State Tracking for Question Generation
- HAL: Improved Text-Image Matching by Mitigating Visual Semantic Hubs
- Toward Filament Segmentation Using Deep Neural Networks
- CMSN: Continuous Multi-stage Network and Variable Margin Cosine Loss for Temporal Action Proposal Generation
- Grouping Capsules Based Different Types
- Localization-aware Channel Pruning for Object Detection
- Distilling Pixel-Wise Feature Similarities for Semantic Segmentation
- Fine-Grained Object Detection over Scientific Document Images with Region Embeddings
- Team PFDet's Methods for Open Images Challenge 2019
- A Locating Model for Pulmonary Tuberculosis Diagnosis in Radiographs
- Gastroscopic Panoramic View: Application to Automatic Polyps Detection under Gastroscopy
- Practical License Plate Recognition in Unconstrained Surveillance Systems with Adversarial Super-Resolution
- IFR-Net: Iterative Feature Refinement Network for Compressed Sensing MRI
- Restoring Spatially-Heterogeneous Distortions using Mixture of Experts Network
- Few-shot Object Detection with Self-adaptive Attention Network for Remote Sensing Images
- Renovating Parsing R-CNN for Accurate Multiple Human Parsing
- Semi-Anchored Detector for One-Stage Object Detection
- Radar+RGB Attentive Fusion for Robust Object Detection in Autonomous Vehicles
- CCA: Exploring the Possibility of Contextual Camouflage Attack on Object Detection
- A Smartphone-based System for Real-time Early Childhood Caries Diagnosis
- What leads to generalization of object proposals?
- Weight Equalizing Shift Scaler-Coupled Post-training Quantization
- Deep Learning-based Human Detection for UAVs with Optical and Infrared Cameras: System and Experiments
- Switching Loss for Generalized Nucleus Detection in Histopathology
- Hard Negative Samples Emphasis Tracker without Anchors
- Multi-Level Temporal Pyramid Network for Action Detection
- Self-supervised Object Tracking with Cycle-consistent Siamese Networks
- U2-ONet: A Two-level Nested Octave U-structure with Multiscale Attention Mechanism for Moving Instances Segmentation
- SADet: Learning An Efficient and Accurate Pedestrian Detector
- AABO: Adaptive Anchor Box Optimization for Object Detection via Bayesian Sub-sampling
- InfoFocus: 3D Object Detection for Autonomous Driving with Dynamic Information Modeling
- Compare and Reweight: Distinctive Image Captioning Using Similar Images Sets
- Understanding Object Detection Through An Adversarial Lens
- Temporal Self-Ensembling Teacher for Semi-Supervised Object Detection
- Domain Contrast for Domain Adaptive Object Detection
- Expandable YOLO: 3D Object Detection from RGB-D Images
- Segmentation task for fashion and apparel
- Improve bone age assessment by learning from anatomical local regions
- Location-Aware Feature Selection Text Detection Network
- DDD20 End-to-End Event Camera Driving Dataset: Fusing Frames and Events with Deep Learning for Improved Steering Prediction
- Large Scale Font Independent Urdu Text Recognition System
- Seismic Shot Gather Noise Localization Using a Multi-Scale Feature-Fusion-Based Neural Network
- GraftNet: An Engineering Implementation of CNN for Fine-grained Multi-label Task
- Detecting and Tracking Communal Bird Roosts in Weather Radar Data
- Farmland Parcel Delineation Using Spatio-temporal Convolutional Networks
- Probabilistic Oriented Object Detection in Automotive Radar
- BiFNet: Bidirectional Fusion Network for Road Segmentation
- Spatial Priming for Detecting Human-Object Interactions
- Semantic Image Manipulation Using Scene Graphs
- Rapid Detection of Aircrafts in Satellite Imagery based on Deep Neural Networks
- A Structure-Aware Relation Network for Thoracic Diseases Detection and Segmentation
- Disentangled Motif-aware Graph Learning for Phrase Grounding
- Human De-occlusion: Invisible Perception and Recovery for Humans
- Decoupled Spatial Temporal Graphs for Generic Visual Grounding
- PatchNet -- Short-range Template Matching for Efficient Video Processing
- Quality-Aware Network for Human Parsing
- Iterative Shrinking for Referring Expression Grounding Using Deep Reinforcement Learning
- Learning from Counting: Leveraging Temporal Classification for Weakly Supervised Object Localization and Detection
- Differentiable Neural Architecture Learning for Efficient Neural Network Design
- GridTracer: Automatic Mapping of Power Grids using Deep Learning and Overhead Imagery
- Balance-Oriented Focal Loss with Linear Scheduling for Anchor Free Object Detection
- PanoNet3D: Combining Semantic and Geometric Understanding for LiDARPoint Cloud Detection
- A Response Retrieval Approach for Dialogue Using a Multi-Attentive Transformer
- Scene Text Detection with Scribble Lines
- SuperOCR: A Conversion from Optical Character Recognition to Image Captioning
- SRF-GAN: Super-Resolved Feature GAN for Multi-Scale Representation
- Selective Spatio-Temporal Aggregation Based Pose Refinement System: Towards Understanding Human Activities in Real-World Videos
- Amadeus: Scalable, Privacy-Preserving Live Video Analytics
- End-to-end Deep Learning Methods for Automated Damage Detection in Extreme Events at Various Scales
- Compositional Scalable Object SLAM
- Multi-layer Feature Aggregation for Deep Scene Parsing Models
- End-to-end trainable network for degraded license plate detection via vehicle-plate relation mining
- Disaster mapping from satellites: damage detection with crowdsourced point labels
- DPNET: Dual-Path Network for Efficient Object Detectioj with Lightweight Self-Attention
- EMDS-7: Environmental Microorganism Image Dataset Seventh Version for Multiple Object Detection Evaluation
- Understanding Egocentric Hand-Object Interactions from Hand Pose Estimation
- Why Do We Click: Visual Impression-aware News Recommendation
- Scene Graph Generation for Better Image Captioning?
- Exploiting Activation based Gradient Output Sparsity to Accelerate Backpropagation in CNNs
- AdaPruner: Adaptive Channel Pruning and Effective Weights Inheritance
- RefineCap: Concept-Aware Refinement for Image Captioning
- An Attention Module for Convolutional Neural Networks
- The Multi-Modal Video Reasoning and Analyzing Competition
- FaPN: Feature-aligned Pyramid Network for Dense Image Prediction
- Exploring Transferable and Robust Adversarial Perturbation Generation from the Perspective of Network Hierarchy
- Disentangle Your Dense Object Detector
- Real-time Keypoints Detection for Autonomous Recovery of the Unmanned Ground Vehicle
- Neighbor-view Enhanced Model for Vision and Language Navigation
- Multi-Modal Pedestrian Detection with Large Misalignment Based on Modal-Wise Regression and Multi-Modal IoU
- Cell Detection from Imperfect Annotation by Pseudo Label Selection Using P-classification
- Automatic and explainable grading of meningiomas from histopathology images
- CFTrack: Center-based Radar and Camera Fusion for 3D Multi-Object Tracking
- Positive-unlabeled Learning for Cell Detection in Histopathology Images with Incomplete Annotations
- Universal Adder Neural Networks
- Humble Teachers Teach Better Students for Semi-Supervised Object Detection
- Discriminative Triad Matching and Reconstruction for Weakly Referring Expression Grounding
- Semi-Autoregressive Transformer for Image Captioning
- FCPose: Fully Convolutional Multi-Person Pose Estimation with Dynamic Instance-Aware Convolutions
- Linguistic Structures as Weak Supervision for Visual Scene Graph Generation
- BCNet: Searching for Network Width with Bilaterally Coupled Network
- FGR: Frustum-Aware Geometric Reasoning for Weakly Supervised 3D Vehicle Detection
- Drill the Cork of Information Bottleneck by Inputting the Most Important Data
- Decoupled IoU Regression for Object Detection
- A Fast Knowledge Distillation Framework for Visual Recognition
- Consensus Graph Representation Learning for Better Grounded Image Captioning
- Semantic-Aware Environment Perception for Mobile Human-Robot Interaction
- Dense Object Detection Based on De-homogenized Queries
- RGB-D Tracking via Hierarchical Modality Aggregation and Distribution Network
- Box2Poly: Memory-Efficient Polygon Prediction of Arbitrarily Shaped and Rotated Text
- The MIS Check-Dam Dataset for Object Detection and Instance Segmentation Tasks
- Contrastive Object-level Pre-training with Spatial Noise Curriculum Learning
- Detection of E-scooter Riders in Naturalistic Scenes
- Head and Body: Unified Detector and Graph Network for Person Search in Media
- Mask Transfiner for High-Quality Instance Segmentation
- Understanding Pixel-level 2D Image Semantics with 3D Keypoint Knowledge Engine
- FooDI-ML: a large multi-language dataset of food, drinks and groceries images and descriptions
- HR-RCNN: Hierarchical Relational Reasoning for Object Detection
- Progressive Hard-case Mining across Pyramid Levels for Object Detection
- A Better Loss for Visual-Textual Grounding
- Neural Rays for Occlusion-aware Image-based Rendering
- Modeling Explicit Concerning States for Reinforcement Learning in Visual Dialogue
- End-to-end Compression Towards Machine Vision: Network Architecture Design and Optimization
- TabLeX: A Benchmark Dataset for Structure and Content Information Extraction from Scientific Tables
- CASIA-Face-Africa: A Large-scale African Face Image Database
- Probabilistic Ranking-Aware Ensembles for Enhanced Object Detections
- RADDet: Range-Azimuth-Doppler based Radar Object Detection for Dynamic Road Users
- MODS -- A USV-oriented object detection and obstacle segmentation benchmark
- Cross-Modal Generative Augmentation for Visual Question Answering
- EagerMOT: 3D Multi-Object Tracking via Sensor Fusion
- Pylot: A Modular Platform for Exploring Latency-Accuracy Tradeoffs in Autonomous Vehicles
- Rethinking Image-Scaling Attacks: The Interplay Between Vulnerabilities in Machine Learning Systems
- EXplainable Neural-Symbolic Learning (X-NeSyL) methodology to fuse deep learning representations with expert knowledge graphs: the MonuMAI cultural heritage use case
- Simple multi-dataset detection
- MosaicOS: A Simple and Effective Use of Object-Centric Images for Long-Tailed Object Detection
- MLPerf Mobile Inference Benchmark
- A Unified Mixture-View Framework for Unsupervised Representation Learning
- Interpretable Visual Reasoning via Induced Symbolic Space
- From Recognition to Prediction: Analysis of Human Action and Trajectory Prediction in Video
- Object Tracking Using Spatio-Temporal Future Prediction
- Learning Spatio-Appearance Memory Network for High-Performance Visual Tracking
- Multi-Label Activity Recognition using Activity-specific Features and Activity Correlations
- Learning Person Re-identification Models from Videos with Weak Supervision
- A Deep Ordinal Distortion Estimation Approach for Distortion Rectification
- AQD: Towards Accurate Fully-Quantized Object Detection
- COBE: Contextualized Object Embeddings from Narrated Instructional Video
- Representation Sharing for Fast Object Detector Search and Beyond
- Deep Learning Methods for Real-time Detection and Analysis of Wagner Ulcer Classification System
- Can 3D Adversarial Logos Cloak Humans?
- Region Proposal Network with Graph Prior and IoU-Balance Loss for Landmark Detection in 3D Ultrasound
- Multi-Plateau Ensemble for Endoscopic Artefact Segmentation and Detection
- On the effectiveness of convolutional autoencoders on image-based personalized recommender systems
- ElixirNet: Relation-aware Network Architecture Adaptation for Medical Lesion Detection
- Face Anti-Spoofing by Learning Polarization Cues in a Real-World Scenario
- The State of Lifelong Learning in Service Robots: Current Bottlenecks in Object Perception and Manipulation
- End-to-end Autonomous Driving Perception with Sequential Latent Representation Learning
- Coupled Network for Robust Pedestrian Detection with Gated Multi-Layer Feature Extraction and Deformable Occlusion Handling
- Training Object Detectors from Few Weakly-Labeled and Many Unlabeled Images
- Oriented Objects as pairs of Middle Lines
- Defective Convolutional Networks
- Large Scale Open-Set Deep Logo Detection
- Deep Mixture Density Network for Probabilistic Object Detection
- To What Extent Does Downsampling, Compression, and Data Scarcity Impact Renal Image Analysis?
- Searching for Accurate Binary Neural Architectures
- Understanding the Effects of Pre-Training for Object Detectors via Eigenspectrum
- Linear Context Transform Block
- Temporal Reasoning Graph for Activity Recognition
- An Objectness Score for Accurate and Fast Detection during Navigation
- High Performance Visual Object Tracking with Unified Convolutional Networks
- Image Captioning with Unseen Objects
- Efficient Method for Categorize Animals in the Wild
- Reprojection R-CNN: A Fast and Accurate Object Detector for 360° Images
- Learning Predicates as Functions to Enable Few-shot Scene Graph Prediction
- Efficient Object Embedding for Spliced Image Retrieval
- Training Quantized Neural Networks with a Full-precision Auxiliary Module
- Geometry-constrained Car Recognition Using a 3D Perspective Network
- MUSCO: Multi-Stage Compression of neural networks
- FVNet: 3D Front-View Proposal Generation for Real-Time Object Detection from Point Clouds
- A Single-shot Object Detector with Feature Aggragation and Enhancement
- Learning More with Less: Conditional PGGAN-based Data Augmentation for Brain Metastases Detection Using Highly-Rough Annotation on MR Images
- DOSED: a deep learning approach to detect multiple sleep micro-events in EEG signal
- A Simple Non-i.i.d. Sampling Approach for Efficient Training and Better Generalization
- Spatially-weighted Anomaly Detection
- Towards a Generic Diver-Following Algorithm: Balancing Robustness and Efficiency in Deep Visual Detection
- Object Detection from Scratch with Deep Supervision
- Weakly- and Semi-Supervised Panoptic Segmentation
- LiDAR and Camera Detection Fusion in a Real Time Industrial Multi-Sensor Collision Avoidance System
- Two at Once: Enhancing Learning and Generalization Capacities via IBN-Net
- Face-Cap: Image Captioning using Facial Expression Analysis
- Deep Global-Connected Net With The Generalized Multi-Piecewise ReLU Activation in Deep Learning
- Object Localization with a Weakly Supervised CapsNet
- Survey of Face Detection on Low-quality Images
- End-to-End Saliency Mapping via Probability Distribution Prediction
- Deep cross-domain building extraction for selective depth estimation from oblique aerial imagery
- Learnable Histogram: Statistical Context Features for Deep Neural Networks
- The Automatic Identification of Butterfly Species
- Visual Manipulation Relationship Network
- Temporally Identity-Aware SSD with Attentional LSTM
- DeepVoting: A Robust and Explainable Deep Network for Semantic Part Detection under Partial Occlusion
- Reformulating Level Sets as Deep Recurrent Neural Network Approach to Semantic Segmentation
- Face Detection with End-to-End Integration of a ConvNet and a 3D Model
- Real-Time Grasping Strategies Using Event Camera
- FAN: Focused Attention Networks
- Exposing Semantic Segmentation Failures via Maximum Discrepancy Competition
- Multi-Kernel Diffusion CNNs for Graph-Based Learning on Point Clouds
- Autonomous detection of molecular configurations in microscopic images based on deep convolutional neural network
- Efficient Model Performance Estimation via Feature Histories
- Self-Supervised Person Detection in 2D Range Data using a Calibrated Camera
- Learning to Navigate for Fine-grained Classification
- Are object detection assessment criteria ready for maritime computer vision?
- Separating Skills and Concepts for Novel Visual Question Answering
- Semi-supervised Cell Detection in Time-lapse Images Using Temporal Consistency
- Advanced Multiple Linear Regression Based Dark Channel Prior Applied on Dehazing Image and Generating Synthetic Haze
- Boosting ship detection in SAR images with complementary pretraining techniques
- Structure Information is the Key: Self-Attention RoI Feature Extractor in 3D Object Detection
- Achieving Human Parity on Visual Question Answering
- Region-Manipulated Fusion Networks for Pancreatitis Recognition
- Sequential End-to-end Network for Efficient Person Search
- UniMoCo: Unsupervised, Semi-Supervised and Full-Supervised Visual Representation Learning
- An Unsupervised Sampling Approach for Image-Sentence Matching Using Document-Level Structural Information
- Deep Lidar CNN to Understand the Dynamics of Moving Vehicles
- Single Shot Scene Text Retrieval
- What Can You Learn from Your Muscles? Learning Visual Representation from Human Interactions
- Co-Grounding Networks with Semantic Attention for Referring Expression Comprehension in Videos
- Stereo Object Matching Network
- Learning Polar Encodings for Arbitrary-Oriented Ship Detection in SAR Images
- Eliminating the Blind Spot: Adapting 3D Object Detection and Monocular Depth Estimation to 360° Panoramic Imagery
- THAT: Two Head Adversarial Training for Improving Robustness at Scale
- From Multi-View to Hollow-3D: Hallucinated Hollow-3D R-CNN for 3D Object Detection
- Defect-GAN: High-Fidelity Defect Synthesis for Automated Defect Inspection
- REGRAD: A Large-Scale Relational Grasp Dataset for Safe and Object-Specific Robotic Grasping in Clutter
- Deformation Robust Roto-Scale-Translation Equivariant CNNs
- Fully Convolutional Scene Graph Generation
- Improving Network Slimming with Nonconvex Regularization
- Semi-Online Knowledge Distillation
- Federated Few-Shot Learning with Adversarial Learning
- Perceptual Attention-based Predictive Control
- DIODE: Dilatable Incremental Object Detection
- Towards Object Detection from Motion
- TL-SDD: A Transfer Learning-Based Method for Surface Defect Detection with Few Samples
- Dense Scene Multiple Object Tracking with Box-Plane Matching
- CAROM -- Vehicle Localization and Traffic Scene Reconstruction from Monocular Cameras on Road Infrastructures
- Localizing Visual Sounds the Hard Way
- Exploiting Multi-Object Relationships for Detecting Adversarial Attacks in Complex Scenes
- Learning Context-Aware Embedding for Person Search
- ZS-SLR: Zero-Shot Sign Language Recognition from RGB-D Videos
- Temporal Knowledge Consistency for Unsupervised Visual Representation Learning
- Learning Visual Affordances with Target-Orientated Deep Q-Network to Grasp Objects by Harnessing Environmental Fixtures
- Fashion-Guided Adversarial Attack on Person Segmentation
- Characters Detection on Namecard with faster RCNN
- MetricOpt: Learning to Optimize Black-Box Evaluation Metrics
- Learning Transferable 3D Adversarial Cloaks for Deep Trained Detectors
- Decoupled Dynamic Filter Networks
- Updatable Siamese Tracker with Two-stage One-shot Learning
- Where are the Blobs: Counting by Localization with Point Supervision
- Multi Voxel-Point Neurons Convolution (MVPConv) for Fast and Accurate 3D Deep Learning
- Decoupled Gradient Harmonized Detector for Partial Annotation: Application to Signet Ring Cell Detection
- Unsupervised Industrial Anomaly Detection via Pattern Generative and Contrastive Networks
- Goal-driven text descriptions for images
- Extract and Merge: Merging extracted humans from different images utilizing Mask R-CNN
- Image-based Detection of Surface Defects in Concrete during Construction
- Query-based Hard-Image Retrieval for Object Detection at Test Time
- Fast Template Matching and Update for Video Object Tracking and Segmentation
- Learning Question-Guided Video Representation for Multi-Turn Video Question Answering
- Lightweight Mask R-CNN for Long-Range Wireless Power Transfer Systems
- Deep Learning-based mitosis detection in breast cancer histologic samples
- RFBTD: RFB Text Detector
- Deep neural networks can be improved using human-derived contextual expectations
- Document Layout Annotation: Database and Benchmark in the Domain of Public Affairs
- ShapeCaptioner: Generative Caption Network for 3D Shapes by Learning a Mapping from Parts Detected in Multiple Views to Sentences
- Multimodal Incremental Transformer with Visual Grounding for Visual Dialogue Generation
- Facial Action Unit Detection on ICU Data for Pain Assessment
- Mask R-CNN Based Object Detection for Intelligent Wireless Power Transfer
- Measuring and Predicting Tag Importance for Image Retrieval
- Multi-Scale Time-Frequency Attention for Acoustic Event Detection
- Hybrid Attention for Automatic Segmentation of Whole Fetal Head in Prenatal Ultrasound Volumes
- Weakly Supervised Dataset Collection for Robust Person Detection
- Towards Embodied Scene Description
- Unsupervised Domain Alignment to Mitigate Low Level Dataset Biases
- TSDM: Tracking by SiamRPN++ with a Depth-refiner and a Mask-generator
- A Multimodal Sentiment Dataset for Video Recommendation
- Scope Head for Accurate Localization in Object Detection
- EDDA: Explanation-driven Data Augmentation to Improve Explanation Faithfulness
- Video-based Person Re-identification without Bells and Whistles
- WixUp: A General Data Augmentation Framework for Wireless Perception in Tracking of Humans
- K-Shot Contrastive Learning of Visual Features with Multiple Instance Augmentations
- Unsupervised data augmentation for object detection
- More Than Just Attention: Improving Cross-Modal Attentions with Contrastive Constraints for Image-Text Matching
- Tensor Composition Net for Visual Relationship Prediction
- You Cannot Easily Catch Me: A Low-Detectable Adversarial Patch for Object Detectors
- Social Behavioral Phenotyping of Drosophila with a2D-3D Hybrid CNN Framework
- mr2NST: Multi-Resolution and Multi-Reference Neural Style Transfer for Mammography
- Weakly Supervised Action Selection Learning in Video
- Detecting Curve Text with Local Segmentation Network and Curve Connection
- MSDU-net: A Multi-Scale Dilated U-net for Blur Detection
- Learning Instance-wise Sparsity for Accelerating Deep Models
- Towards Robust Pattern Recognition: A Review
- Semantic Dense Reconstruction with Consistent Scene Segments
- Exploiting Visual Semantic Reasoning for Video-Text Retrieval
- GPR: Grasp Pose Refinement Network for Cluttered Scenes
- AE TextSpotter: Learning Visual and Linguistic Representation for Ambiguous Text Spotting
- Proposal Flow: Semantic Correspondences from Object Proposals
- Transformer Assisted Convolutional Network for Cell Instance Segmentation
- Materials In Paintings (MIP): An interdisciplinary dataset for perception, art history, and computer vision
- Combating Ambiguity for Hash-code Learning in Medical Instance Retrieval
- Optical deep learning nano-profilometry
- Horizontal-to-Vertical Video Conversion
- Deep Learning Acceleration Techniques for Real Time Mobile Vision Applications
- Image Captioning with Compositional Neural Module Networks
- Saliency deep embedding for aurora image search
- AutoTrajectory: Label-free Trajectory Extraction and Prediction from Videos using Dynamic Points
- Location-Aware Box Reasoning for Anchor-Based Single-Shot Object Detection
- Robust and Efficient Graph Correspondence Transfer for Person Re-identification
- Are you doing what I say? On modalities alignment in ALFRED
- Slot Based Image Augmentation System for Object Detection
- Understanding of Emotion Perception from Art
- Semi-Autoregressive Image Captioning
- Domain2Vec: Domain Embedding for Unsupervised Domain Adaptation
- 2nd Place Solution to ECCV 2020 VIPriors Object Detection Challenge
- Deep Learning Estimation of Absorbed Dose for Nuclear Medicine Diagnostics
- ALET (Automated Labeling of Equipment and Tools): A Dataset, a Baseline and a Usecase for Tool Detection in the Wild
- LGA-RCNN: Loss-Guided Attention for Object Detection
- DDR-ID: Dual Deep Reconstruction Networks Based Image Decomposition for Anomaly Detection
- MT: Multi-Perspective Feature Learning Network for Scene Text Detection
- LSTC: Boosting Atomic Action Detection with Long-Short-Term Context
- Adma: A Flexible Loss Function for Neural Networks
- Understanding Video Content: Efficient Hero Detection and Recognition for the Game "Honor of Kings"
- Improving Object Detection with Selective Self-supervised Self-training
- Learning Instance-Aware Object Detection Using Determinantal Point Processes
- Commonality-Parsing Network across Shape and Appearance for Partially Supervised Instance Segmentation
- Locality-Aware Rotated Ship Detection in High-Resolution Remote Sensing Imagery Based on Multi-Scale Convolutional Network
- Osteoporosis Prescreening using Panoramic Radiographs through a Deep Convolutional Neural Network with Attention Mechanism
- Person Identification with Visual Summary for a Safe Access to a Smart Home
- Prime-Aware Adaptive Distillation
- Learning with Rethinking: Recurrently Improving Convolutional Neural Networks through Feedback
- Human Motion Capture Using a Drone
- Multi-Stream Attention Learning for Monocular Vehicle Velocity and Inter-Vehicle Distance Estimation
- Camera On-boarding for Person Re-identification using Hypothesis Transfer Learning
- FPAN: Fine-grained and Progressive Attention Localization Network for Data Retrieval
- Enabling Incremental Knowledge Transfer for Object Detection at the Edge
- Clouds of Oriented Gradients for 3D Detection of Objects, Surfaces, and Indoor Scene Layouts
- Search Spaces for Neural Model Training
- A Joint Intensity-Neuromorphic Event Imaging System for Resource Constrained Devices
- Relevance Attack on Detectors
- Removing Word-Level Spurious Alignment between Images and Pseudo-Captions in Unsupervised Image Captioning
- Learning Relation Alignment for Calibrated Cross-modal Retrieval
- Training a Binary Weight Object Detector by Knowledge Transfer for Autonomous Driving
- Feedback Attention for Cell Image Segmentation
- Certainty Driven Consistency Loss on Multi-Teacher Networks for Semi-Supervised Learning
- MOS: A Low Latency and Lightweight Framework for Face Detection, Landmark Localization, and Head Pose Estimation
- Dynamic Resolution Network
- On Applying Machine Learning/Object Detection Models for Analysing Digitally Captured Physical Prototypes from Engineering Design Projects
- Object-Aware Multi-Branch Relation Networks for Spatio-Temporal Video Grounding
- SeqDialN: Sequential Visual Dialog Networks in Joint Visual-Linguistic Representation Space
- QuadricSLAM: Dual Quadrics from Object Detections as Landmarks in Object-oriented SLAM
- Strawberry Detection using Mixed Training on Simulated and Real Data
- DORi: Discovering Object Relationship for Moment Localization of a Natural-Language Query in Video
- Temporal Action Localization with Variance-Aware Networks
- Few-Shot Object Detection via Knowledge Transfer
- Open-World Entity Segmentation
- Deep Volumetric Universal Lesion Detection using Light-Weight Pseudo 3D Convolution and Surface Point Regression
- Image Captioning with Integrated Bottom-Up and Multi-level Residual Top-Down Attention for Game Scene Understanding
- Task-driven Semantic Coding via Reinforcement Learning
- LabelEnc: A New Intermediate Supervision Method for Object Detection
- supervised adptive threshold network for instance segmentation
- Beyond the Deep Metric Learning: Enhance the Cross-Modal Matching with Adversarial Discriminative Domain Regularization
- Financial ticket intelligent recognition system based on deep learning
- Robust Regression via Deep Negative Correlation Learning
- 3D Object Recognition By Corresponding and Quantizing Neural 3D Scene Representations
- Few-shot Learning with Global Relatedness Decoupled-Distillation
- Greedy Network Enlarging
- Visible Feature Guidance for Crowd Pedestrian Detection
- AFAN: Augmented Feature Alignment Network for Cross-Domain Object Detection
- Curriculum Learning with Diversity for Supervised Computer Vision Tasks
- Multiple interaction learning with question-type prior knowledge for constraining answer search space in visual question answering
- Unsupervised Object Detection with LiDAR Clues
- Progressive Stage-wise Learning for Unsupervised Feature Representation Enhancement
- Localize to Classify and Classify to Localize: Mutual Guidance in Object Detection
- Dynamic Graph: Learning Instance-aware Connectivity for Neural Networks
- An Empirical Study of DNNs Robustification Inefficacy in Protecting Visual Recommenders
- FDNAS: Improving Data Privacy and Model Diversity in AutoML
- ULSD: Unified Line Segment Detection across Pinhole, Fisheye, and Spherical Cameras
- STELA: A Real-Time Scene Text Detector with Learned Anchor
- An Action Recognition network for specific target based on rMC and RPN
- Rethinking Task and Metrics of Instance Segmentation on 3D Point Clouds
- ASSD: Attentive Single Shot Multibox Detector
- Multi-Expert Adversarial Attack Detection in Person Re-identification Using Context Inconsistency
- Multi-Head Attention with Diversity for Learning Grounded Multilingual Multimodal Representations
- LIP: Learning Instance Propagation for Video Object Segmentation
- A Self Validation Network for Object-Level Human Attention Estimation
- Analysis and a Solution of Momentarily Missed Detection for Anchor-based Object Detectors
- Single View Physical Distance Estimation using Human Pose
- Enforcing Reasoning in Visual Commonsense Reasoning
- An Autonomous Approach to Measure Social Distances and Hygienic Practices during COVID-19 Pandemic in Public Open Spaces
- Learning Temporal Action Proposals With Fewer Labels
- NADS-Net: A Nimble Architecture for Driver and Seat Belt Detection via Convolutional Neural Networks
- ActBERT: Learning Global-Local Video-Text Representations
- Bi-Dimensional Feature Alignment for Cross-Domain Object Detection
- Group-based Distinctive Image Captioning with Memory Attention
- Progressive Unsupervised Person Re-identification by Tracklet Association with Spatio-Temporal Regularization
- Discriminative Semantic Feature Pyramid Network with Guided Anchoring for Logo Detection
- Event Recognition with Automatic Album Detection based on Sequential Processing, Neural Attention and Image Captioning
- Temporal Action Localization using Long Short-Term Dependency
- The Devil is in the Boundary: Exploiting Boundary Representation for Basis-based Instance Segmentation
- Towards Real-Time Advancement of Underwater Visual Quality with GAN
- Congestion Analysis of Convolutional Neural Network-Based Pedestrian Counting Methods on Helicopter Footage
- Enabling Highly Efficient Capsule Networks Processing Through A PIM-Based Architecture Design
- Is First Person Vision Challenging for Object Tracking?
- Energy Drain of the Object Detection Processing Pipeline for Mobile Devices: Analysis and Implications
- Single-Shot Multi-Person 3D Pose Estimation From Monocular RGB
- Location-aware Upsampling for Semantic Segmentation
- Ellipse Detection and Localization with Applications to Knots in Sawn Lumber Images
- A gamified simulator and physical platform for self-driving algorithm training and validation
- Descriptor-Free Multi-View Region Matching for Instance-Wise 3D Reconstruction
- Batch Normalization with Enhanced Linear Transformation
- Learning a metacognition for object perception
- SISA: Securing Images by Selective Alteration
- Shape Prior Non-Uniform Sampling Guided Real-time Stereo 3D Object Detection
- Context-Gated Convolution
- Weak Supervision helps Emergence of Word-Object Alignment and improves Vision-Language Tasks
- Global Context Aware RCNN for Object Detection
- DMRM: A Dual-channel Multi-hop Reasoning Model for Visual Dialog
- Maria: A Visual Experience Powered Conversational Agent
- Visual Agreement Regularized Training for Multi-Modal Machine Translation
- DeepACEv2: Automated Chromosome Enumeration in Metaphase Cell Images Using Deep Convolutional Neural Networks
- Aggressive Perception-Aware Navigation using Deep Optical Flow Dynamics and PixelMPC
- ParaNet: Deep Regular Representation for 3D Point Clouds
- Frequency Disentangled Residual Network
- Deep Residual Dense U-Net for Resolution Enhancement in Accelerated MRI Acquisition
- Efficient Human Pose Estimation with Depthwise Separable Convolution and Person Centroid Guided Joint Grouping
- Joint Distribution Alignment via Adversarial Learning for Domain Adaptive Object Detection
- MixTConv: Mixed Temporal Convolutional Kernels for Efficient Action Recogntion
- Deep Learning based Multi-Modal Sensing for Tracking and State Extraction of Small Quadcopters
- ShuffleDet: Real-Time Vehicle Detection Network in On-board Embedded UAV Imagery
- One-Shot Domain Adaptation For Face Generation
- Object Instance Mining for Weakly Supervised Object Detection
- Robust Segmentation of Optic Disc and Cup from Fundus Images Using Deep Neural Networks
- Inverting and Understanding Object Detectors
- Widening and Squeezing: Towards Accurate and Efficient QNNs
- Answer-checking in Context: A Multi-modal FullyAttention Network for Visual Question Answering
- A Holistically-Guided Decoder for Deep Representation Learning with Applications to Semantic Segmentation and Object Detection
- Real-time Human-Robot Collaborative Manipulations of Cylindrical and Cubic Objects via Geometric Primitives and Depth Information
- AIBench: An Agile Domain-specific Benchmarking Methodology and an AI Benchmark Suite
- EHSOD: CAM-Guided End-to-end Hybrid-Supervised Object Detection with Cascade Refinement
- Object 6D Pose Estimation with Non-local Attention
- Feature-Driven Super-Resolution for Object Detection
- Global Context Networks
- Expressing Objects just like Words: Recurrent Visual Embedding for Image-Text Matching
- Towards Understanding the Effectiveness of Attention Mechanism
- Colonoscopy Polyp Detection: Domain Adaptation From Medical Report Images to Real-time Videos
- An Artificial Intelligence System for Combined Fruit Detection and Georeferencing, Using RTK-Based Perspective Projection in Drone Imagery
- Morphological Networks for Image De-raining
- PVSS: A Progressive Vehicle Search System for Video Surveillance Networks
- Hierarchy Denoising Recursive Autoencoders for 3D Scene Layout Prediction
- SOLO: A Simple Framework for Instance Segmentation
- PandaNet : Anchor-Based Single-Shot Multi-Person 3D Pose Estimation
- Towards Fully Automated Manga Translation
- Colorectal Polyp Detection in Real-world Scenario: Design and Experiment Study
- Towards Pedestrian Detection Using RetinaNet in ECCV 2018 Wider Pedestrian Detection Challenge
- WDR FACE: The First Database for Studying Face Detection in Wide Dynamic Range
- Simultaneous x, y Pixel Estimation and Feature Extraction for Multiple Small Objects in a Scene: A Description of the ALIEN Network
- Box-level Segmentation Supervised Deep Neural Networks for Accurate and Real-time Multispectral Pedestrian Detection
- Min-Entropy Latent Model for Weakly Supervised Object Detection
- Mixup Without Hesitation
- LiDAR-assisted Large-scale Privacy Protection in Street-view Cycloramas
- You Only Look & Listen Once: Towards Fast and Accurate Visual Grounding
- Towards Cross-Modal Forgery Detection and Localization on Live Surveillance Videos
- LIT: Light-field Inference of Transparency for Refractive Object Localization
- SurveilEdge: Real-time Video Query based on Collaborative Cloud-Edge Deep Learning
- Towards Deep Learning Assisted Autonomous UAVs for Manipulation Tasks in GPS-Denied Environments
- Deep Reinforcement Learning for Active Human Pose Estimation
- Natural Language Video Localization with Learnable Moment Proposals
- Improved Techniques for Quantizing Deep Networks with Adaptive Bit-Widths
- Real Time Incremental Foveal Texture Mapping for Autonomous Vehicles
- An Improvement of Object Detection Performance using Multi-step Machine Learnings
- Biologically Inspired Visual System Architecture for Object Recognition in Autonomous Systems
- Weak Novel Categories without Tears: A Survey on Weak-Shot Learning
- Mask-GD Segmentation Based Robotic Grasp Detection
- Scene-Aware Error Modeling of LiDAR/Visual Odometry for Fusion-based Vehicle Localization
- Global Convergence and Geometric Characterization of Slow to Fast Weight Evolution in Neural Network Training for Classifying Linearly Non-Separable Data
- L-Verse: Bidirectional Generation Between Image and Text
- Progressive Cross-Stream Cooperation in Spatial and Temporal Domain for Action Localization
- R-TOD: Real-Time Object Detector with Minimized End-to-End Delay for Autonomous Driving
- LANCE: Efficient Low-Precision Quantized Winograd Convolution for Neural Networks Based on Graphics Processing Units
- Semantic Segmentation for Compound figures
- Visual Question Answering based on Local-Scene-Aware Referring Expression Generation
- Recognizing Object Affordances to Support Scene Reasoning for Manipulation Tasks
- Introducing Pose Consistency and Warp-Alignment for Self-Supervised 6D Object Pose Estimation in Color Images
- Vehicle Detection in Deep Learning
- MLMA-Net: multi-level multi-attentional learning for multi-label object detection in textile defect images
- MLOD: A multi-view 3D object detection based on robust feature fusion method
- PointINS: Point-based Instance Segmentation
- An End-to-End Food Image Analysis System
- Domain Adaptation Regularization for Spectral Pruning
- Histogram of Cell Types: Deep Learning for Automated Bone Marrow Cytology
- Denoising Large-Scale Image Captioning from Alt-text Data using Content Selection Models
- BiDet: An Efficient Binarized Object Detector
- Zero-Shot Scene Graph Relation Prediction through Commonsense Knowledge Integration
- PBRnet: Pyramidal Bounding Box Refinement to Improve Object Localization Accuracy
- IC Networks: Remodeling the Basic Unit for Convolutional Neural Networks
- Bringing in the outliers: A sparse subspace clustering approach to learn a dictionary of mouse ultrasonic vocalizations
- Train a One-Million-Way Instance Classifier for Unsupervised Visual Representation Learning
- DEPARA: Deep Attribution Graph for Deep Knowledge Transferability
- Natural Adversarial Objects
- Neural Twins Talk
- Multi-Person Pose Estimation with Enhanced Feature Aggregation and Selection
- Quantifying the presence of graffiti in urban environments
- A New Few-shot Segmentation Network Based on Class Representation
- BABO: Background Activation Black-Out for Efficient Object Detection
- Pseudo-Labeling for Small Lesion Detection on Diabetic Retinopathy Images
- Neural Mesh Refiner for 6-DoF Pose Estimation
- Live Reconstruction of Large-Scale Dynamic Outdoor Worlds
- Single Shot Video Object Detector
- Impression Space from Deep Template Network
- Tackling the Problem of Limited Data and Annotations in Semantic Segmentation
- Learning to Parse Wireframes in Images of Man-Made Environments
- ContourRend: A Segmentation Method for Improving Contours by Rendering
- Tracking the Untrackable
- Graph-Based Social Relation Reasoning
- Volumetric Transformer Networks
- SparseTrain: Exploiting Dataflow Sparsity for Efficient Convolutional Neural Networks Training
- Visual Relation Grounding in Videos
- Computer-aided Tumor Diagnosis in Automated Breast Ultrasound using 3D Detection Network
- LiDAR point-cloud processing based on projection methods: a comparison
- Implicit Saliency in Deep Neural Networks
- Hierarchical Scene Parsing by Weakly Supervised Learning with Image Descriptions
- Leveraging Localization for Multi-camera Association
- Low-Light Maritime Image Enhancement with Regularized Illumination Optimization and Deep Noise Suppression
- Woodpecker-DL: Accelerating Deep Neural Networks via Hardware-Aware Multifaceted Optimizations
- Sparse Coding Driven Deep Decision Tree Ensembles for Nuclear Segmentation in Digital Pathology Images
- Using Cross-Model EgoSupervision to Learn Cooperative Basketball Intention
- AntiDote: Attention-based Dynamic Optimization for Neural Network Runtime Efficiency
- Layout-induced Video Representation for Recognizing Agent-in-Place Actions
- Object Detection in the Context of Mobile Augmented Reality
- CFAD: Coarse-to-Fine Action Detector for Spatiotemporal Action Localization
- Vision-based deep execution monitoring
- Bosch Deep Learning Hardware Benchmark
- Graphical Object Detection in Document Images
- Fast Single-shot Ship Instance Segmentation Based on Polar Template Mask in Remote Sensing Images
- Learning Fixation Point Strategy for Object Detection and Classification
- Online Spatiotemporal Action Detection and Prediction via Causal Representations
- Modification method for single-stage object detectors that allows to exploit the temporal behaviour of a scene to improve detection accuracy
- Few-shot Object Detection with Feature Attention Highlight Module in Remote Sensing Images
- Geometric Neural Phrase Pooling: Modeling the Spatial Co-occurrence of Neurons
- Tasks Integrated Networks: Joint Detection and Retrieval for Image Search
- An unsupervised deep learning framework via integrated optimization of representation learning and GMM-based modeling
- Learning Joint Representations of Videos and Sentences with Web Image Search
- Learning semantic Image attributes using Image recognition and knowledge graph embeddings
- Deep Collective Learning: Learning Optimal Inputs and Weights Jointly in Deep Neural Networks
- CoFF: Cooperative Spatial Feature Fusion for 3D Object Detection on Autonomous Vehicles
- Self-grouping Convolutional Neural Networks
- Smart-Inspect: Micro Scale Localization and Classification of Smartphone Glass Defects for Industrial Automation
- Addressing Class Imbalance in Scene Graph Parsing by Learning to Contrast and Score
- Deep Local Global Refinement Network for Stent Analysis in IVOCT Images
- Feature Fusion Detector for Semantic Cognition of Remote Sensing
- Hidden State Guidance: Improving Image Captioning using An Image Conditioned Autoencoder
- Regularizing Neural Networks via Stochastic Branch Layers
- Semi-Automatic Crowdsourcing Tool for Online Food Image Collection and Annotation
- Organ At Risk Segmentation with Multiple Modality
- Assisting human experts in the interpretation of their visual process: A case study on assessing copper surface adhesive potency
- Literature Review: Human Segmentation with Static Camera
- ROI Pooled Correlation Filters for Visual Tracking
- Investigations of the Influences of a CNN's Receptive Field on Segmentation of Subnuclei of Bilateral Amygdalae
- A Computing Kernel for Network Binarization on PyTorch
- Mini Lesions Detection on Diabetic Retinopathy Images via Large Scale CNN Features
- Computer-Aided Clinical Skin Disease Diagnosis Using CNN and Object Detection Models
- Adaptive Multi-scale Detection of Acoustic Events
- Occluded Pedestrian Detection with Visible IoU and Box Sign Predictor
- OptiBox: Breaking the Limits of Proposals for Visual Grounding
- Better Understanding Hierarchical Visual Relationship for Image Caption
- Water Supply Prediction Based on Initialized Attention Residual Network
- Autonomous Removal of Perspective Distortion for Robotic Elevator Button Recognition
- A Multi-oriented Chinese Keyword Spotter Guided by Text Line Detection
- A Holistic Approach for Data-Driven Object Cutout
- Multi-Scale Weight Sharing Network for Image Recognition
- Contextual Sense Making by Fusing Scene Classification, Detections, and Events in Full Motion Video
- IG-TRACK: IOU Guided Siamese Networks for visual object tracking
- SafeNet: An Assistive Solution to Assess Incoming Threats for Premises
- Unsupervised Domain Adaptive Object Detection using Forward-Backward Cyclic Adaptation
- Measuring the Utilization of Public Open Spaces by Deep Learning: a Benchmark Study at the Detroit Riverfront
- Joint Deep Learning of Facial Expression Synthesis and Recognition
- Learning Hyperspectral Feature Extraction and Classification with ResNeXt Network
- Rotational Rectification Network: Enabling Pedestrian Detection for Mobile Vision
- Layered Embeddings for Amodal Instance Segmentation
- An End-to-End Framework for Unsupervised Pose Estimation of Occluded Pedestrians
- Deep CTR Prediction in Display Advertising
- Learning Object Scale With Click Supervision for Object Detection
- Joint 2D-3D Breast Cancer Classification
- A Video Analysis Method on Wanfang Dataset via Deep Neural Network
- Weeping and Gnashing of Teeth: Teaching Deep Learning in Image and Video Processing Classes
- Residual-CNDS for Grand Challenge Scene Dataset
- Effect of Various Regularizers on Model Complexities of Neural Networks in Presence of Input Noise
- Vision Based Picking System for Automatic Express Package Dispatching
- Dataset Culling: Towards Efficient Training Of Distillation-Based Domain Specific Models
- Detecting Gaze Towards Eyes in Natural Social Interactions and its Use in Child Assessment
- Segmentation of Objects by Hashing
- Spiking Neural Network based Region Proposal Networks for Neuromorphic Vision Sensors
- IvaNet: Learning to jointly detect and segment objets with the help of Local Top-Down Modules
- Weakly Supervised Instance Segmentation by Deep Community Learning
- On Class Imbalance and Background Filtering in Visual Relationship Detection
- CycleCluster: Modernising Clustering Regularisation for Deep Semi-Supervised Classification
- Gun Source and Muzzle Head Detection
- Weighted Average Precision: Adversarial Example Detection in the Visual Perception of Autonomous Vehicles
- 3D Object Detection on Point Clouds using Local Ground-aware and Adaptive Representation of scenes' surface
- Competing Ratio Loss for Discriminative Multi-class Image Classification
- Network Pruning via Annealing and Direct Sparsity Control
- Localizing Interpretable Multi-scale informative Patches Derived from Media Classification Task
- Two-stage multi-scale breast mass segmentation for full mammogram analysis without user intervention
- On evaluating CNN representations for low resource medical image classification
- A High-Performance Object Proposals based on Horizontal High Frequency Signal
- OVC-Net: Object-Oriented Video Captioning with Temporal Graph and Detail Enhancement
- Learning Object Permanence from Video
- TapLab: A Fast Framework for Semantic Video Segmentation Tapping into Compressed-Domain Knowledge
- Pose Augmentation: Class-agnostic Object Pose Transformation for Object Recognition
- Using constraint structure and an improved object detection network to detect the 12^{th} Vertebra from CT images with a limited field of view for image-guided radiotherapy
- iFAN: Image-Instance Full Alignment Networks for Adaptive Object Detection
- Convolutional Neural Networks for User Identificationbased on Motion Sensors Represented as Image
- Pacemaker: Intermediate Teacher Knowledge Distillation For On-The-Fly Convolutional Neural Network
- Interaction Graphs for Object Importance Estimation in On-road Driving Videos
- ZSTAD: Zero-Shot Temporal Activity Detection
- Affinity Graph Supervision for Visual Recognition
- Towards Detection of Sheep Onboard a UAV
- DMV: Visual Object Tracking via Part-level Dense Memory and Voting-based Retrieval
- Self-Guided Adaptation: Progressive Representation Alignment for Domain Adaptive Object Detection
- On the Evaluation of Prohibited Item Classification and Detection in Volumetric 3D Computed Tomography Baggage Security Screening Imagery
- Scalable Change Retrieval Using Deep 3D Neural Codes
- Conditionally Learn to Pay Attention for Sequential Visual Task
- Long Short-Term Relation Networks for Video Action Detection
- EOLO: Embedded Object Segmentation only Look Once
- Machine Vision for Improved Human-Robot Cooperation in Adverse Underwater Conditions
- Spatio-temporal Tubelet Feature Aggregation and Object Linking in Videos
- Visual Question Answering Using Semantic Information from Image Descriptions
- Deep Learning Meets SAR
- Learning to Reweight with Deep Interactions
- A method for detecting text of arbitrary shapes in natural scenes that improves text spotting
- Learning Camera-Aware Noise Models
- Enhancing and Learning Denoiser without Clean Reference
- Group Whitening: Balancing Learning Efficiency and Representational Capacity
- Feature Flow: In-network Feature Flow Estimation for Video Object Detection
- Graph-based Heuristic Search for Module Selection Procedure in Neural Module Network
- Self-Guided Multiple Instance Learning for Weakly Supervised Disease Classification and Localization in Chest Radiographs
- Can images help recognize entities? A study of the role of images for Multimodal NER
- Artificial Intelligence: Research Impact on Key Industries; the Upper-Rhine Artificial Intelligence Symposium (UR-AI 2020)
- Neighbourhood Distillation: On the benefits of non end-to-end distillation
- Large Product Key Memory for Pretrained Language Models
- Contralaterally Enhanced Networks for Thoracic Disease Detection
- Interpretable Neural Computation for Real-World Compositional Visual Question Answering
- WeightAlign: Normalizing Activations by Weight Alignment
- Adaptive and Azimuth-Aware Fusion Network of Multimodal Local Features for 3D Object Detection
- Learning to Reconstruct and Segment 3D Objects
- Image Captioning with Visual Object Representations Grounded in the Textual Modality
- Universal Bounding Box Regression and Its Applications
- Learning Quantized Neural Nets by Coarse Gradient Method for Non-linear Classification
- Lattice Fusion Networks for Image Denoising
- Resolving the cybersecurity Data Sharing Paradox to scale up cybersecurity via a co-production approach towards data sharing
- Deep Learning for Recognizing Mobile Targets in Satellite Imagery
- Self-supervised asymmetric deep hashing with margin-scalable constraint
- Deep Point-wise Prediction for Action Temporal Proposal
- Recognition of Russian traffic signs in winter conditions. Solutions of the "Ice Vision" competition winners
- OBoW: Online Bag-of-Visual-Words Generation for Self-Supervised Learning
- Weakly Supervised Localization Using Background Images
- Gaussian Temporal Awareness Networks for Action Localization
- Semi-Automatic Annotation For Visual Object Tracking
- Self-Adaptive Reconfigurable Arrays (SARA): Using ML to Assist Scaling GEMM Acceleration
- A DCNN-based Arbitrarily-Oriented Object Detector for Quality Control and Inspection Application
- POD: Practical Object Detection with Scale-Sensitive Network
- Hands-on Guidance for Distilling Object Detectors
- Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation
- Decoupled Box Proposal and Featurization with Ultrafine-Grained Semantic Labels Improve Image Captioning and Visual Question Answering
- GPNAS: A Neural Network Architecture Search Framework Based on Graphical Predictor
- Scene-based Factored Attention for Image Captioning
- Generation and Simulation of Yeast Microscopy Imagery with Deep Learning
- HishabNet: Detection, Localization and Calculation of Handwritten Bengali Mathematical Expressions
- Characterizing and Improving the Robustness of Self-Supervised Learning through Background Augmentations
- Adaptive Binary-Ternary Quantization
- Revisiting the Loss Weight Adjustment in Object Detection
- USB: Universal-Scale Object Detection Benchmark
- Efficient Online Transfer Learning for 3D Object Classification in Autonomous Driving
- Universal Physical Camouflage Attacks on Object Detectors
- Effect of Visual Extensions on Natural Language Understanding in Vision-and-Language Models
- Training Deep Neural Networks via Branch-and-Bound
- Robotic Waste Sorter with Agile Manipulation and Quickly Trainable Detector
- Opening up Open-World Tracking
- VGNMN: Video-grounded Neural Module Network to Video-Grounded Language Tasks
- Few-Shot Video Object Detection
- Image Segmentation, Compression and Reconstruction from Edge Distribution Estimation with Random Field and Random Cluster Theories
- Plot and Rework: Modeling Storylines for Visual Storytelling
- FishNet: A Camera Localizer using Deep Recurrent Networks
- CUAB: Convolutional Uncertainty Attention Block Enhanced the Chest X-ray Image Analysis
- Human Object Interaction Detection using Two-Direction Spatial Enhancement and Exclusive Object Prior
- T-EMDE: Sketching-based global similarity for cross-modal retrieval
- Instance-aware Remote Sensing Image Captioning with Cross-hierarchy Attention
- Rethinking of Radar's Role: A Camera-Radar Dataset and Systematic Annotator via Coordinate Alignment
- Wearable Travel Aid for Environment Perception and Navigation of Visually Impaired People
- Distilling Image Classifiers in Object Detectors
- Place recognition survey: An update on deep learning approaches
- Goal-oriented Object Importance Estimation in On-road Driving Videos
- Class agnostic moving target detection by color and location prediction of moving area
- Deep Bregman Divergence for Contrastive Learning of Visual Representations
- Partially-Supervised Novel Object Captioning Leveraging Context from Paired Data
- StereOBJ-1M: Large-scale Stereo Image Dataset for 6D Object Pose Estimation
- GoG: Relation-aware Graph-over-Graph Network for Visual Dialog
- ISyNet: Convolutional Neural Networks design for AI accelerator
- LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
- Recognizing Part Attributes with Insufficient Data
- Sequential Voting with Relational Box Fields for Active Object Detection
- Knowledge-driven Active Learning
- SwAMP: Swapped Assignment of Multi-Modal Pairs for Cross-Modal Retrieval
- TraVLR: Now You See It, Now You Don't! A Bimodal Dataset for Evaluating Visio-Linguistic Reasoning
- Multi-Glimpse Network: A Robust and Efficient Classification Architecture based on Recurrent Downsampled Attention
- A Survey on Green Deep Learning
- Graph Relation Transformer: Incorporating pairwise object features into the Transformer architecture
- Unsupervised Spiking Instance Segmentation on Event Data using STDP
- Searching for TrioNet: Combining Convolution with Local and Global Self-Attention
- Visual-Relation Conscious Image Generation from Structured-Text
- Evaluating Adversarial Attacks on ImageNet: A Reality Check on Misclassification Classes
- ScarfNet: Multi-scale Features with Deeply Fused and Redistributed Semantics for Enhanced Object Detection
- Realistic simulation of users for IT systems in cyber ranges
- Automated Detection of Patients in Hospital Video Recordings
- NoFADE: Analyzing Diminishing Returns on CO2 Investment
- Attention-based Feature Decomposition-Reconstruction Network for Scene Text Detection
- Flood Analytics Information System (FAIS) Version 4.00 Manual
- Semi-Supervised Surface Anomaly Detection of Composite Wind Turbine Blades From Drone Imagery
- Human-Object Interaction Detection via Weak Supervision
- Object-Centric Unsupervised Image Captioning
- Object Detection in Specific Traffic Scenes using YOLOv2
- Can neural networks count digit frequency?
- Predictive Ensemble Learning with Application to Scene Text Detection
- Open-set 3D Object Detection
- Vision Pair Learning: An Efficient Training Framework for Image Classification
- Document Dewarping with Control Points
- Implicit Label Augmentation on Partially Annotated Clips via Temporally-Adaptive Features Learning
- A Wireless-Vision Dataset for Privacy Preserving Human Activity Recognition
- MVB: A Large-Scale Dataset for Baggage Re-Identification and Merged Siamese Networks
- Efficient Transfer Learning via Joint Adaptation of Network Architecture and Weight
- Attention Filtering for Multi-person Spatiotemporal Action Detection on Deep Two-Stream CNN Architectures
- Towards a Sample Efficient Reinforcement Learning Pipeline for Vision Based Robotics
- Soccer Player Tracking in Low Quality Video
- DTNN: Energy-efficient Inference with Dendrite Tree Inspired Neural Networks for Edge Vision Applications
- Issues in Object Detection in Videos using Common Single-Image CNNs
- PSRR-MaxpoolNMS: Pyramid Shifted MaxpoolNMS with Relationship Recovery
- Dual Normalization Multitasking for Audio-Visual Sounding Object Localization
- Pathology-Aware Generative Adversarial Networks for Medical Image Augmentation
- PDPGD: Primal-Dual Proximal Gradient Descent Adversarial Attack
- Transferable Adversarial Examples for Anchor Free Object Detection
- Not Only Look But Observe: Variational Observation Model of Scene-Level 3D Multi-Object Understanding for Probabilistic SLAM
- Understanding top-down attention using task-oriented ablation design
- Labeled Data Generation with Inexact Supervision
- Implicit Feature Alignment: Learn to Convert Text Recognizer to Text Spotter
- ShuffleBlock: Shuffle to Regularize Deep Convolutional Neural Networks
- Region Refinement Network for Salient Object Detection
- Quantized Neural Networks via {-1, +1} Encoding Decomposition and Acceleration
- A Framework for Real-time Traffic Trajectory Tracking, Speed Estimation, and Driver Behavior Calibration at Urban Intersections Using Virtual Traffic Lanes
- Generation and frame characteristics of predefined evenly-distributed class centroids for pattern classification
- SRPN: similarity-based region proposal networks for nuclei and cells detection in histology images
- Building Intelligent Autonomous Navigation Agents
- Real-time 3D Object Detection using Feature Map Flow
- Learning from Web Data with Self-Organizing Memory Module
- Adventurer's Treasure Hunt: A Transparent System for Visually Grounded Compositional Visual Question Answering based on Scene Graphs
- SRF-Net: Selective Receptive Field Network for Anchor-Free Temporal Action Detection
- Doing good by fighting fraud: Ethical anti-fraud systems for mobile payments
- Integration of Text-maps in Convolutional Neural Networks for Region Detection among Different Textual Categories
- Unsupervised Image Segmentation by Mutual Information Maximization and Adversarial Regularization
- Semantics to Space(S2S): Embedding semantics into spatial space for zero-shot verb-object query inferencing
- A Transductive Maximum Margin Classifier for Few-Shot Learning
- Automated Object Behavioral Feature Extraction for Potential Risk Analysis based on Video Sensor
- Integrating Deep Learning and Augmented Reality to Enhance Situational Awareness in Firefighting Environments
- Case Relation Transformer: A Crossmodal Language Generation Model for Fetching Instructions
- CompConv: A Compact Convolution Module for Efficient Feature Learning
- Merging Tasks for Video Panoptic Segmentation
- Doctor of Crosswise: Reducing Over-parametrization in Neural Networks
- CS-R-FCN: Cross-supervised Learning for Large-Scale Object Detection
- Bidirectional Regression for Arbitrary-Shaped Text Detection
- Semantically-Aware Strategies for Stereo-Visual Robotic Obstacle Avoidance
- FSD: Feature Skyscraper Detector for Stem End and Blossom End of Navel Orange
- Multi-Modal Temporal Convolutional Network for Anticipating Actions in Egocentric Videos
- Geometric Data Augmentation Based on Feature Map Ensemble
- Multi-stage Pre-training over Simplified Multimodal Pre-training Models
- United We Learn Better: Harvesting Learning Improvements From Class Hierarchies Across Tasks
- Domain Adaptor Networks for Hyperspectral Image Recognition
- Double-Dot Network for Antipodal Grasp Detection
- Neural Twins Talk & Alternative Calculations
- Two is a crowd: tracking relations in videos
- Identifying and Exploiting Structures for Reliable Deep Learning
- Hardware realisation of nonlinear dynamical systems for and from biology
- Leveraging Orientation for Weakly Supervised Object Detection with Application to Firearm Localization
- Exploring Data Aggregation and Transformations to Generalize across Visual Domains
- OOWL500: Overcoming Dataset Collection Bias in the Wild
- Improving Object Detection and Attribute Recognition by Feature Entanglement Reduction
- Target-Tailored Source-Transformation for Scene Graph Generation
- Product-oriented Machine Translation with Cross-modal Cross-lingual Pre-training
- Layer-wise Customized Weak Segmentation Block and AIoU Loss for Accurate Object Detection
- PR Product: A Substitute for Inner Product in Neural Networks
- Densely Semantic Enhancement for Domain Adaptive Region-free Detectors
- AIP: Adversarial Iterative Pruning Based on Knowledge Transfer for Convolutional Neural Networks
- Architecture Aware Latency Constrained Sparse Neural Networks
- Vision-and-Language or Vision-for-Language? On Cross-Modal Influence in Multimodal Transformers
- Temporal RoI Align for Video Object Recognition
- Multi-Scale Aligned Distillation for Low-Resolution Detection
- Fast query-by-example speech search using separable model
- Learning Natural Language Generation from Scratch
- COVR: A test-bed for Visually Grounded Compositional Generalization with real images
- MOMENTA: A Multimodal Framework for Detecting Harmful Memes and Their Targets
- EllipseNet: Anchor-Free Ellipse Detection for Automatic Cardiac Biometrics in Fetal Echocardiography
- -Cal: Calibrated aleatoric uncertainty estimation from neural networks for robot perception
- Turning old models fashion again: Recycling classical CNN networks using the Lattice Transformation
- Adversarial Semantic Contour for Object Detection
- CrossCLR: Cross-modal Contrastive Learning For Multi-modal Video Representations
- Deep Learning Approach Protecting Privacy in Camera-Based Critical Applications
- CNN-based Human Detection for UAVs in Search and Rescue
- Learning Single/Multi-Attribute of Object with Symmetry and Group
- Multi-Precision Quantized Neural Networks via Encoding Decomposition of -1 and +1
- Learning Deep Image Priors for Blind Image Denoising
- Topic Scene Graph Generation by Attention Distillation from Caption
- Towards Mixed-Precision Quantization of Neural Networks via Constrained Optimization
- Student Helping Teacher: Teacher Evolution via Self-Knowledge Distillation
- ASK: Adaptively Selecting Key Local Features for RGB-D Scene Recognition
- Self-Supervised Object Detection via Generative Image Synthesis
- Mitosis Detection for Breast Cancer Pathology Images Using UV-Net
- NAS-FCOS: Efficient Search for Object Detection Architectures
- CoopSubNet: Cooperating Subnetwork for Data-Driven Regularization of Deep Networks under Limited Training Budgets
- Generative Residual Attention Network for Disease Detection
- AlteregoNets: a way to human augmentation
- Continuous Trade-off Optimization between Fast and Accurate Deep Face Detectors
- CBIR using Pre-Trained Neural Networks
- PolyTrack: Tracking with Bounding Polygons
- Learned Image Compression for Machine Perception
- Bootstrap Your Object Detector via Mixed Training
- Frustum Fusion: Pseudo-LiDAR and LiDAR Fusion for 3D Detection
- Net: Augmented Parallel-Pyramid Net for Attention Guided Pose Estimation
- Fast Local Attack: Generating Local Adversarial Examples for Object Detectors
- Spatial-temporal Fusion Convolutional Neural Network for Simulated Driving Behavior Recognition
- Adversarial Domain Randomization
- Efficient Pipelines for Vision-Based Context Sensing
- DSIC: Dynamic Sample-Individualized Connector for Multi-Scale Object Detection
- Transformation Driven Visual Reasoning
- Fast Object Detection with Latticed Multi-Scale Feature Fusion
- Stereo Frustums: A Siamese Pipeline for 3D Object Detection
- Automated data extraction of bar chart raster images
- GRCNN: Graph Recognition Convolutional Neural Network for Synthesizing Programs from Flow Charts
- Data-efficient Alignment of Multimodal Sequences by Aligning Gradient Updates and Internal Feature Distributions
- FPAENet: Pneumonia Detection Network Based on Feature Pyramid Attention Enhancement
- Context Encoding Chest X-rays
- Unsupervised Part Discovery via Feature Alignment
- A Hypergradient Approach to Robust Regression without Correspondence
- Towards Part-Based Understanding of RGB-D Scans
- Copyspace: Where to Write on Images?
- Discovering Spatio-Temporal Action Tubes
- PCGAN: Partition-Controlled Human Image Generation
- Learning to Generate Content-Aware Dynamic Detectors
- DeepKey: Towards End-to-End Physical Key Replication From a Single Photograph
- Robust Real-Time Pedestrian Detection on Embedded Devices
- Smoothed Gaussian Mixture Models for Video Classification and Recommendation
- Dense Multiscale Feature Fusion Pyramid Networks for Object Detection in UAV-Captured Images
- Impoved RPN for Single Targets Detection based on the Anchor Mask Net
- AmphibianDetector: adaptive computation for moving objects detection
- Research on Fast Text Recognition Method for Financial Ticket Image
- Fine-Grained Vehicle Perception via 3D Part-Guided Visual Data Augmentation
- Adaptive Remote Sensing Image Attribute Learning for Active Object Detection
- Shape Back-Projection In 3D Scenes
- Semi Supervised Deep Quick Instance Detection and Segmentation
- Multi-Task Network Pruning and Embedded Optimization for Real-time Deployment in ADAS
- Online Active Proposal Set Generation for Weakly Supervised Object Detection
- Hierarchical Graph-RNNs for Action Detection of Multiple Activities
- Discovering Multi-Label Actor-Action Association in a Weakly Supervised Setting
- Evolution Attack On Neural Networks
- Provident Vehicle Detection at Night: The PVDN Dataset
- A Unified Light Framework for Real-time Fault Detection of Freight Train Images
- Cross Chest Graph for Disease Diagnosis with Structural Relational Reasoning
- Multi-Attribute Enhancement Network for Person Search
- Custom Object Detection via Multi-Camera Self-Supervised Learning
- APEX-Net: Automatic Plot Extractor Network
- Analysing object detectors from the perspective of co-occurring object categories
- MASON: A Model AgnoStic ObjectNess Framework
- Anchor Distance for 3D Multi-Object Distance Estimation from 2D Single Shot
- Towards Large-Scale Video Video Object Mining
- OICSR: Out-In-Channel Sparsity Regularization for Compact Deep Neural Networks
- Phase Space Reconstruction Network for Lane Intrusion Action Recognition
- Relationship-based Neural Baby Talk
- Robust 2D/3D Vehicle Parsing in CVIS
- Sim-to-Real Transfer of Robot Learning with Variable Length Inputs
- Conceptual Text Region Network: Cognition-Inspired Accurate Scene Text Detection
- The Case for High-Accuracy Classification: Think Small, Think Many!
- Neural Task Planning with And-Or Graph Representations
- Structure-Aware Face Clustering on a Large-Scale Graph with Nodes
- Picasso: A CUDA-based Library for Deep Learning over 3D Meshes
- A Multiplexed Network for End-to-End, Multilingual OCR
- Knowledge Distillation By Sparse Representation Matching
- Joint Learning of Neural Transfer and Architecture Adaptation for Image Recognition
- Two-phase weakly supervised object detection with pseudo ground truth mining
- Malignancy Prediction and Lesion Identification from Clinical Dermatological Images
- Beyond Short Clips: End-to-End Video-Level Learning with Collaborative Memories
- Fingerspelling Detection in American Sign Language
- Deep Online Fused Video Stabilization
- Visual Tracking via Reliable Memories
- TraMNet - Transition Matrix Network for Efficient Action Tube Proposals
- Few-Shot Transformation of Common Actions into Time and Space
- Learning Spatial Context with Graph Neural Network for Multi-Person Pose Grouping
- AsymmNet: Towards ultralight convolution neural networks using asymmetrical bottlenecks
- Weakly Supervised Object Localization and Detection: A Survey
- Error-Corrected Margin-Based Deep Cross-Modal Hashing for Facial Image Retrieval
- LU-Net: a multi-task network to improve the robustness of segmentation of left ventriclular structures by deep learning in 2D echocardiography
- End-to-End Learning via a Convolutional Neural Network for Cancer Cell Line Classification
- MA 3 : Model Agnostic Adversarial Augmentation for Few Shot learning
- Training few-shot classification via the perspective of minibatch and pretraining
- Feature Lenses: Plug-and-play Neural Modules for Transformation-Invariant Visual Representations
- Intelligent Orchestration of ADAS Pipelines on Next Generation Automotive Platforms
- Spatially Attentive Output Layer for Image Classification
- CPARR: Category-based Proposal Analysis for Referring Relationships
- MultiPoseNet: Fast Multi-Person Pose Estimation using Pose Residual Network
- Generative Adversarial Networks for Unsupervised Object Co-localization
- Instance Segmentation of Biomedical Images with an Object-aware Embedding Learned with Local Constraints
- A Revised Generative Evaluation of Visual Dialogue
- Long Activity Video Understanding using Functional Object-Oriented Network
- Localizing Grouped Instances for Efficient Detection in Low-Resource Scenarios
- Cross-Modality Relevance for Reasoning on Language and Vision
- Anchors Based Method for Fingertips Position Estimation from a Monocular RGB Image using Deep Neural Network
- Ventral-Dorsal Neural Networks: Object Detection via Selective Attention
- A Deep Learning-based Radar and Camera Sensor Fusion Architecture for Object Detection
- Synthesizing Unrestricted False Positive Adversarial Objects Using Generative Models
- KL-Divergence-Based Region Proposal Network for Object Detection
- Improving 3D Object Detection for Pedestrians with Virtual Multi-View Synthesis Orientation Estimation
- DeepFirearm: Learning Discriminative Feature Representation for Fine-grained Firearm Retrieval
- The Research of the Real-time Detection and Recognition of Targets in Streetscape Videos
- Delving into the Imbalance of Positive Proposals in Two-stage Object Detection
- Accelerating Neural Network Inference by Overflow Aware Quantization
- Condensing Two-stage Detection with Automatic Object Key Part Discovery
- End-to-end Learning for Inter-Vehicle Distance and Relative Velocity Estimation in ADAS with a Monocular Camera
- Physics-based Scene-level Reasoning for Object Pose Estimation in Clutter
- Cascaded Regression Tracking: Towards Online Hard Distractor Discrimination
- Sequential Feature Filtering Classifier
- Off-Policy Self-Critical Training for Transformer in Visual Paragraph Generation
- Orthogonal Deep Models As Defense Against Black-Box Attacks
- Motion Prediction in Visual Object Tracking
- A Few-Shot Sequential Approach for Object Counting
- A Classification approach towards Unsupervised Learning of Visual Representations