Gaussian Error Linear Units (GELUs)
arXiv:1606.08415
Abstract
We propose the Gaussian Error Linear Unit (GELU), a high-performing neural network activation function. The GELU activation function is , where the standard Gaussian cumulative distribution function. The GELU nonlinearity weights inputs by their value, rather than gates inputs by their sign as in ReLUs (). We perform an empirical evaluation of the GELU nonlinearity against the ReLU and ELU activations and find performance improvements across all considered computer vision, natural language processing, and speech tasks.
Trimmed version of 2016 draft
References in corpus (6)
- SGDR: Stochastic Gradient Descent with Warm Restarts
- Learning with Pseudo-Ensembles
- Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations
- Deep Residual Networks with Exponential Linear Unit
- Natural Neural Networks
- Adjusting for Dropout Variance in Batch Normalization and Weight Initialization
Cited by in corpus (433)
- Learning Transferable Visual Models From Natural Language Supervision
- EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
- A Survey on Visual Transformer
- Attention Mechanisms in Computer Vision: A Survey
- PVT v2: Improved Baselines with Pyramid Vision Transformer
- E(3)-Equivariant Graph Neural Networks for Data-Efficient and Accurate Interatomic Potentials
- MLP-Mixer: An all-MLP Architecture for Vision
- Vision Transformers for Single Image Dehazing
- Transformer in Transformer
- ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
- Unified Language Model Pre-training for Natural Language Understanding and Generation
- CoAtNet: Marrying Convolution and Attention for All Data Sizes
- A novel time-frequency Transformer based on self-attention mechanism and its application in fault diagnosis of rolling bearings
- EfficientDet: Scalable and Efficient Object Detection
- Learnable latent embeddings for joint behavioral and neural analysis
- SpookyNet: Learning Force Fields with Electronic Degrees of Freedom and Nonlocal Effects
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
- Distributional Soft Actor-Critic: Off-Policy Reinforcement Learning for Addressing Value Estimation Errors
- Stiff-PINN: Physics-Informed Neural Network for Stiff Chemical Kinetics
- P2T: Pyramid Pooling Transformer for Scene Understanding
- MetaFormer Baselines for Vision
- Action Transformer: A Self-Attention Model for Short-Time Pose-Based Human Action Recognition
- nnFormer: Interleaved Transformer for Volumetric Segmentation
- Transformer-based Acoustic Modeling for Hybrid Speech Recognition
- High-Performance Large-Scale Image Recognition Without Normalization
- The Principles of Deep Learning Theory
- Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation
- Decoding speech perception from non-invasive brain recordings
- Dual Cross-Attention for Medical Image Segmentation
- Perceiver IO: A General Architecture for Structured Inputs & Outputs
- You Only Look at One Sequence: Rethinking Transformer in Vision through Object Detection
- Restormer: Efficient Transformer for High-Resolution Image Restoration
- FCN-Transformer Feature Fusion for Polyp Segmentation
- Language Modeling with Deep Transformers
- ST-GRAT: A Novel Spatio-temporal Graph Attention Network for Accurately Forecasting Dynamically Changing Road Speed
- CTCNet: A CNN-Transformer Cooperation Network for Face Image Super-Resolution
- Keyword Transformer: A Self-Attention Model for Keyword Spotting
- Uncovering the Limits of Adversarial Training against Norm-Bounded Adversarial Examples
- Physics-guided deep learning framework for predictive modeling of the Reynolds stress anisotropy
- CMU-Net: A Strong ConvMixer-based Medical Ultrasound Image Segmentation Network
- ERNIE: Enhanced Language Representation with Informative Entities
- Perceiver: General Perception with Iterative Attention
- Reliable extrapolation of deep neural operators informed by physics or sparse observations
- SAINT: Improved Neural Networks for Tabular Data via Row Attention and Contrastive Pre-Training
- MLIC: Multi-Reference Entropy Model for Learned Image Compression
- Physics-Informed Neural Networks for Shell Structures
- Uformer: A General U-Shaped Transformer for Image Restoration
- Fixing Data Augmentation to Improve Adversarial Robustness
- When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations
- What Are Bayesian Neural Network Posteriors Really Like?
- Mixout: Effective Regularization to Finetune Large-scale Pretrained Language Models
- Aligning AI With Shared Human Values
- Fine-tuning wav2vec2 for speaker recognition
- CycleMLP: A MLP-like Architecture for Dense Prediction
- Pretrained Transformers as Universal Computation Engines
- StructBERT: Incorporating Language Structures into Pre-training for Deep Language Understanding
- Smooth Adversarial Training
- Strategies for the Construction of Machine-Learning Potentials for Accurate and Efficient Atomic-Scale Simulations
- Transformer-based Spatial-Temporal Feature Learning for EEG Decoding
- VTGAN: Semi-supervised Retinal Image Synthesis and Disease Prediction using Vision Transformers
- A Critical Review of Physics-Informed Machine Learning Applications in Subsurface Energy Systems
- CharBERT: Character-aware Pre-trained Language Model
- Discovering Parametric Activation Functions
- Compacter: Efficient Low-Rank Hypercomplex Adapter Layers
- Ground state energy functional with Hartree-Fock efficiency and chemical accuracy
- HTR-VT: Handwritten Text Recognition with Vision Transformer
- Transformers in Single Object Tracking: An Experimental Survey
- Fourier Neural Operator Surrogate Model to Predict 3D Seismic Waves Propagation
- Benchmarking Detection Transfer Learning with Vision Transformers
- Dive into Deep Learning
- COOT: Cooperative Hierarchical Transformer for Video-Text Representation Learning
- Learning Dynamic Graph Representation of Brain Connectome with Spatio-Temporal Attention
- Hybrid Spectrogram and Waveform Source Separation
- HiCMAE: Hierarchical Contrastive Masked Autoencoder for Self-Supervised Audio-Visual Emotion Recognition
- EfficientPose: An efficient, accurate and scalable end-to-end 6D multi object pose estimation approach
- IA-RED: Interpretability-Aware Redundancy Reduction for Vision Transformers
- Actionable and Interpretable Fault Localization for Recurring Failures in Online Service Systems
- Large Kernel Distillation Network for Efficient Single Image Super-Resolution
- Evolution TANN and the identification of internal variables and evolution equations in solid mechanics
- KPGT: Knowledge-Guided Pre-training of Graph Transformer for Molecular Property Prediction
- Music Demixing Challenge 2021
- Combining Deep Reinforcement Learning and Search for Imperfect-Information Games
- 6D-ViT: Category-Level 6D Object Pose Estimation via Transformer-based Instance Representation Learning
- Guided Depth Map Super-resolution: A Survey
- LRT: An Efficient Low-Light Restoration Transformer for Dark Light Field Images
- CMT: Convolutional Neural Networks Meet Vision Transformers
- T-former: An Efficient Transformer for Image Inpainting
- RDP-Net: Region Detail Preserving Network for Change Detection
- Accurate emulator for the redshift-space power spectrum of dark matter halos and its application to galaxy power spectrum
- Multiscale Vision Transformers
- An Image is Worth 16x16 Words, What is a Video Worth?
- Self-attending RNN for Speech Enhancement to Improve Cross-corpus Generalization
- Deep Learning Methods for Partial Differential Equations and Related Parameter Identification Problems
- Evolving Normalization-Activation Layers
- ReduNet: A White-box Deep Network from the Principle of Maximizing Rate Reduction
- Depth-Wise Convolutions in Vision Transformers for Efficient Training on Small Datasets
- EmoNeXt: an Adapted ConvNeXt for Facial Emotion Recognition
- U-TILISE: A Sequence-to-sequence Model for Cloud Removal in Optical Satellite Time Series
- Graph-MLP: Node Classification without Message Passing in Graph
- Applications of physics informed neural operators
- "No, to the Right" -- Online Language Corrections for Robotic Manipulation via Shared Autonomy
- A neural operator-based surrogate solver for free-form electromagnetic inverse design
- Differentially Private Fine-tuning of Language Models
- Rapid Seismic Waveform Modeling and Inversion with Neural Operators
- Transformer Encoder with Multiscale Deep Learning for Pain Classification Using Physiological Signals
- DeFT-AN: Dense Frequency-Time Attentive Network for Multichannel Speech Enhancement
- Very Deep VAEs Generalize Autoregressive Models and Can Outperform Them on Images
- Towards an astronomical foundation model for stars with a Transformer-based model
- PPT Fusion: Pyramid Patch Transformerfor a Case Study in Image Fusion
- Only Train Once: A One-Shot Neural Network Training And Pruning Framework
- Mask and Reason: Pre-Training Knowledge Graph Transformers for Complex Logical Queries
- ObjFormer: Learning Land-Cover Changes From Paired OSM Data and Optical High-Resolution Imagery via Object-Guided Transformer
- A Modified Batch Intrinsic Plasticity Method for Pre-training the Random Coefficients of Extreme Learning Machines
- The Cosmic Graph: Optimal Information Extraction from Large-Scale Structure using Catalogues
- Fast and Sample-Efficient Interatomic Neural Network Potentials for Molecules and Materials Based on Gaussian Moments
- Evolutionary Preference Learning via Graph Nested GRU ODE for Session-based Recommendation
- Clinically-Inspired Multi-Agent Transformers for Disease Trajectory Forecasting from Multimodal Data
- Trex: Learning Execution Semantics from Micro-Traces for Binary Similarity
- Economic Topology Optimization of District Heating Networks using a Pipe Penalization Approach
- Parameter Efficient Multimodal Transformers for Video Representation Learning
- RiemannONets: Interpretable Neural Operators for Riemann Problems
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data
- Machine learning the derivative discontinuity of density-functional theory
- An equivariant graph neural network for the elasticity tensors of all seven crystal systems
- CT-Net: Arbitrary-Shaped Text Detection via Contour Transformer
- LUKE: Deep Contextualized Entity Representations with Entity-aware Self-attention
- SpectralFormer: Rethinking Hyperspectral Image Classification with Transformers
- MINN: Learning the dynamics of differential-algebraic equations and application to battery modeling
- Concurrent ischemic lesion age estimation and segmentation of CT brain using a Transformer-based network
- Pay Attention to MLPs
- Applications of Scientific Machine Learning for the Analysis of Functionally Graded Porous Beams
- MetaFormer Is Actually What You Need for Vision
- Multimodal Model with Text and Drug Embeddings for Adverse Drug Reaction Classification
- User Retention-oriented Recommendation with Decision Transformer
- Hierarchical Vision Transformers for Cardiac Ejection Fraction Estimation
- S-MLP: Spatial-Shift MLP Architecture for Vision
- FitVid: Overfitting in Pixel-Level Video Prediction
- Thermally Averaged Magnetic Anisotropy Tensors via Machine Learning Based on Gaussian Moments
- SVFAP: Self-supervised Video Facial Affect Perceiver
- Scaling Vision with Sparse Mixture of Experts
- NormFormer: Improved Transformer Pretraining with Extra Normalization
- Visual Relationship Detection with Visual-Linguistic Knowledge from Multimodal Representations
- On the Benefit of Width for Neural Networks: Disappearance of Bad Basins
- Conditional Motion In-betweening
- Pixel super-resolved virtual staining of label-free tissue using diffusion models
- Large language models for automated scholarly paper review: A survey
- Pretrained Transformers Improve Out-of-Distribution Robustness
- Simple Training Strategies and Model Scaling for Object Detection
- How fine can fine-tuning be? Learning efficient language models
- SE(3)-equivariant prediction of molecular wavefunctions and electronic densities
- Mapping Dark Matter in the Milky Way using Normalizing Flows and Gaia DR3
- Interaction-Aware Trajectory Planning for Autonomous Vehicles with Analytic Integration of Neural Networks into Model Predictive Control
- IGLUE: A Benchmark for Transfer Learning across Modalities, Tasks, and Languages
- SRC-Net: Bi-Temporal Spatial Relationship Concerned Network for Change Detection
- Emerging Cross-lingual Structure in Pretrained Language Models
- ConvMLP: Hierarchical Convolutional MLPs for Vision
- X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal Transformers
- Knowledge Fusion and Semantic Knowledge Ranking for Open Domain Question Answering
- Learn molecular representations from large-scale unlabeled molecules for drug discovery
- Accelerating Simulation of Stiff Nonlinear Systems using Continuous-Time Echo State Networks
- Multi-head or Single-head? An Empirical Comparison for Transformer Training
- MCTNet: A Multi-Scale CNN-Transformer Network for Change Detection in Optical Remote Sensing Images
- Fast Dynamic 1D Simulation of Divertor Plasmas with Neural PDE Surrogates
- What Formal Languages Can Transformers Express? A Survey
- Characterizing signal propagation to close the performance gap in unnormalized ResNets
- Magnetohydrodynamics with Physics Informed Neural Operators
- Bolt: Bridging the Gap between Auto-tuners and Hardware-native Performance
- Self-Attentive Hawkes Processes
- PALM: Pre-training an Autoencoding&Autoregressive Language Model for Context-conditioned Generation
- Feature Learning in Infinite-Width Neural Networks
- KG-BART: Knowledge Graph-Augmented BART for Generative Commonsense Reasoning
- SF2Former: Amyotrophic Lateral Sclerosis Identification From Multi-center MRI Data Using Spatial and Frequency Fusion Transformer
- Random Features for Kernel Approximation: A Survey on Algorithms, Theory, and Beyond
- Comparing Machine Learning and Interpolation Methods for Loop-Level Calculations
- HRPVT: High-Resolution Pyramid Vision Transformer for medium and small-scale human pose estimation
- TransEM:Residual Swin-Transformer based regularized PET image reconstruction
- Enhancing Few-shot Image Classification with Cosine Transformer
- Associating Objects with Transformers for Video Object Segmentation
- DeTiME: Diffusion-Enhanced Topic Modeling using Encoder-decoder based LLM
- Visual Parser: Representing Part-whole Hierarchies with Transformers
- Automated Novelty Evaluation of Academic Paper: A Collaborative Approach Integrating Human and Large Language Model Knowledge
- Music Source Separation Based on a Lightweight Deep Learning Framework (DTTNET: DUAL-PATH TFC-TDF UNET)
- Mixed Precision Low-bit Quantization of Neural Network Language Models for Speech Recognition
- ASiT: Local-Global Audio Spectrogram vIsion Transformer for Event Classification
- A Lorentz-Equivariant Transformer for All of the LHC
- Efficient Transformers with Dynamic Token Pooling
- Towards Crowdsourced Training of Large Neural Networks using Decentralized Mixture-of-Experts
- Revisiting Tensor Basis Neural Networks for Reynolds stress modeling: application to plane channel and square duct flows
- Transformer-Driven Inverse Problem Transform for Fast Blind Hyperspectral Image Dehazing
- Pruning Self-attentions into Convolutional Layers in Single Path
- Machine learning modeling of the atomic structure and physical properties of alkali and alkaline-earth aluminosilicate glasses and melts
- GNNBuilder: An Automated Framework for Generic Graph Neural Network Accelerator Generation, Simulation, and Optimization
- AutoFormer: Searching Transformers for Visual Recognition
- I-BERT: Integer-only BERT Quantization
- OnDev-LCT: On-Device Lightweight Convolutional Transformers towards federated learning
- VARA-TTS: Non-Autoregressive Text-to-Speech Synthesis based on Very Deep VAE with Residual Attention
- On Explaining Your Explanations of BERT: An Empirical Study with Sequence Classification
- Continuous PDE Dynamics Forecasting with Implicit Neural Representations
- Hybrid Auxiliary Field Quantum Monte Carlo for Molecular Systems
- Rethinking Channel Dimensions for Efficient Model Design
- MusiCoder: A Universal Music-Acoustic Encoder Based on Transformers
- AReLU: Attention-based Rectified Linear Unit
- INet: Inter-Intra-slice Interpolation Network for Medical Slice Synthesis
- Improving Generalization for Multimodal Fake News Detection
- GRAM: An Interpretable Approach for Graph Anomaly Detection using Gradient Attention Maps
- RikiNet: Reading Wikipedia Pages for Natural Question Answering
- NeuBTF: Neural fields for BTF encoding and transfer
- Categorical Normalizing Flows via Continuous Transformations
- Dual-path Self-Attention RNN for Real-Time Speech Enhancement
- Galaxy clustering from the bottom up: A Streaming Model emulator I
- A Parameter-Efficient Learning Approach to Arabic Dialect Identification with Pre-Trained General-Purpose Speech Model
- Searching for Efficient Multi-Stage Vision Transformers
- NILMFormer: Non-Intrusive Load Monitoring that Accounts for Non-Stationarity
- Zero-touch Continuous Network Slicing Control via Scalable Actor-Critic Learning
- On the locality of local neural operator in learning fluid dynamics
- MM-ALT: A Multimodal Automatic Lyric Transcription System
- Fixed-Dimensional and Permutation Invariant State Representation of Autonomous Driving
- Visual Transformers for Primates Classification and Covid Detection
- Local neural operator for solving transient partial differential equations on varied domains
- Speech Separation with Pretrained Frontend to Minimize Domain Mismatch
- Reverse Ordering Techniques for Attention-Based Channel Prediction
- Thank you for Attention: A survey on Attention-based Artificial Neural Networks for Automatic Speech Recognition
- Counterfactual Data Augmentation using Locally Factored Dynamics
- Knowledge Enhanced Attention for Robust Natural Language Inference
- The Low-Rank Simplicity Bias in Deep Networks
- Forensic Analysis and Localization of Multiply Compressed MP3 Audio Using Transformers
- MLP Singer: Towards Rapid Parallel Korean Singing Voice Synthesis
- A Multidimensional Graph Fourier Transformation Neural Network for Vehicle Trajectory Prediction
- LLIC: Large Receptive Field Transform Coding with Adaptive Weights for Learned Image Compression
- Augmenting Decompiler Output with Learned Variable Names and Types
- A Multi-grained based Attention Network for Semi-supervised Sound Event Detection
- Code-switched inspired losses for generic spoken dialog representations
- Normalized Attention Without Probability Cage
- DirectMultiStep: Direct Route Generation for Multistep Retrosynthesis
- Can weight sharing outperform random architecture search? An investigation with TuNAS
- Eformer: Edge Enhancement based Transformer for Medical Image Denoising
- MixerGAN: An MLP-Based Architecture for Unpaired Image-to-Image Translation
- BayesFT: Bayesian Optimization for Fault Tolerant Neural Network Architecture
- Deep calibration of the quadratic rough Heston model
- A Fast and Robust BERT-based Dialogue State Tracker for Schema-Guided Dialogue Dataset
- Deep neural networks approach to microbial colony detection -- a comparative analysis
- Exploring Continuous Integrate-and-Fire for Adaptive Simultaneous Speech Translation
- CATFace: Cross-Attribute-Guided Transformer with Self-Attention Distillation for Low-Quality Face Recognition
- Generative Pre-Training for Speech with Autoregressive Predictive Coding
- TransMed: Transformers Advance Multi-modal Medical Image Classification
- High-dimensional reinforcement learning for optimization and control of ultracold quantum gases
- Variance-reduced Language Pretraining via a Mask Proposal Network
- Language Models are Good Translators
- Activation Function Optimization Scheme for Image Classification
- A Continuous Convolutional Trainable Filter for Modelling Unstructured Data
- Training speaker recognition systems with limited data
- NAS-VAD: Neural Architecture Search for Voice Activity Detection
- Biologically Inspired Oscillating Activation Functions Can Bridge the Performance Gap between Biological and Artificial Neurons
- Reconstructing nodal pressures in water distribution systems with graph neural networks
- Deep Learning in Diabetic Foot Ulcers Detection: A Comprehensive Evaluation
- GoonDAE: Denoising-Based Driver Assistance for Off-Road Teleoperation
- Complexity Measures for Neural Networks with General Activation Functions Using Path-based Norms
- Equivariant Neural Networks for Spin Dynamics Simulations of Itinerant Magnets
- SAPAG: A Self-Adaptive Privacy Attack From Gradients
- ScaleVLAD: Improving Multimodal Sentiment Analysis via Multi-Scale Fusion of Locally Descriptors
- Is Supervised Syntactic Parsing Beneficial for Language Understanding? An Empirical Investigation
- One4all User Representation for Recommender Systems in E-commerce
- A Novel Stochastic Transformer-based Approach for Post-Traumatic Stress Disorder Detection using Audio Recording of Clinical Interviews
- NEU: A Meta-Algorithm for Universal UAP-Invariant Feature Representation
- Lightweight Convolutional Representations for On-Device Natural Language Processing
- Deep Learning is Singular, and That's Good
- Smooth activations and reproducibility in deep networks
- Multichannel Orthogonal Transform-Based Perceptron Layers for Efficient ResNets
- Shifted Chunk Transformer for Spatio-Temporal Representational Learning
- Quizbowl: The Case for Incremental Question Answering
- Computational limits to the legibility of the imaged human brain
- A Qualitative Study of the Dynamic Behavior for Adaptive Gradient Algorithms
- Generative Model for Constructing Reaction Path from Initial to Final States
- Efficient Image Pre-Training with Siamese Cropped Masked Autoencoders
- Rethinking Implicit Neural Representations for Vision Learners
- Psi-GAN: A power-spectrum-informed generative adversarial network for the emulation of large-scale structure maps across cosmologies and redshifts
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
- CoPR: Towards Accurate Visual Localization With Continuous Place-descriptor Regression
- Optimizing Inference Performance of Transformers on CPUs
- Hire-MLP: Vision MLP via Hierarchical Rearrangement
- Trial by FIRE: Probing the dark matter density profile of dwarf galaxies with GraphNPE
- DAGN: Discourse-Aware Graph Network for Logical Reasoning
- Variable Name Recovery in Decompiled Binary Code using Constrained Masked Language Modeling
- Deep Representation Learning for Open Vocabulary Electroencephalography-to-Text Decoding
- Predictive Modeling: BIM Command Recommendation Based on Large-scale Usage Logs
- Dealing with training and test segmentation mismatch: FBK@IWSLT2021
- Multi-Task Learning for Conversational Question Answering over a Large-Scale Knowledge Base
- LaRa: Latents and Rays for Multi-Camera Bird's-Eye-View Semantic Segmentation
- ADF & TransApp: A Transformer-Based Framework for Appliance Detection Using Smart Meter Consumption Series
- A Neural Network Perturbation Theory Based on the Born Series
- RGB-D-Fusion: Image Conditioned Depth Diffusion of Humanoid Subjects
- Embodied BERT: A Transformer Model for Embodied, Language-guided Visual Task Completion
- Tractable structured natural gradient descent using local parameterizations
- One Self-Configurable Model to Solve Many Abstract Visual Reasoning Problems
- Link Prediction on N-ary Relational Facts: A Graph-based Approach
- Supervising the Transfer of Reasoning Patterns in VQA
- Bilingual Alignment Pre-Training for Zero-Shot Cross-Lingual Transfer
- Sens-BERT: Enabling Transferability and Re-calibration of Calibration Models for Low-cost Sensors under Reference Measurements Scarcity
- Encoding Distributional Soft Actor-Critic for Autonomous Driving in Multi-lane Scenarios
- Self-supervised Learning of Rotation-invariant 3D Point Set Features using Transformer and its Self-distillation
- Bayesian sparsification for deep neural networks with Bayesian model reduction
- Time-Space Transformers for Video Panoptic Segmentation
- IDIAPers @ Causal News Corpus 2022: Efficient Causal Relation Identification Through a Prompt-based Few-shot Approach
- On the Universality of the Double Descent Peak in Ridgeless Regression
- MOI-Mixer: Improving MLP-Mixer with Multi Order Interactions in Sequential Recommendation
- Probabilistic Software Modeling: A Data-driven Paradigm for Software Analysis
- Exoplanet Transit Candidate Identification in TESS Full-Frame Images via a Transformer-Based Algorithm
- Symmetrical Gaussian Error Linear Units (SGELUs)
- Beyond Point Estimate: Inferring Ensemble Prediction Variation from Neuron Activation Strength in Recommender Systems
- Language Modelling for Source Code with Transformer-XL
- Our Evaluation Metric Needs an Update to Encourage Generalization
- Mesa: A Memory-saving Training Framework for Transformers
- Self-Attention Gazetteer Embeddings for Named-Entity Recognition
- Extract then Distill: Efficient and Effective Task-Agnostic BERT Distillation
- BD-MSA: Body decouple VHR Remote Sensing Image Change Detection method guided by multi-scale feature information aggregation
- Improving Robustness using Generated Data
- Unified Multi-modal Diagnostic Framework with Reconstruction Pre-training and Heterogeneity-combat Tuning
- OkwuGbé: End-to-End Speech Recognition for Fon and Igbo
- Boundary Aware U-Net for Glacier Segmentation
- Enriching Non-Autoregressive Transformer with Syntactic and SemanticStructures for Neural Machine Translation
- SGD-QA: Fast Schema-Guided Dialogue State Tracking for Unseen Services
- Learning to Retrieve Entity-Aware Knowledge and Generate Responses with Copy Mechanism for Task-Oriented Dialogue Systems
- Know Your Limits: Uncertainty Estimation with ReLU Classifiers Fails at Reliable OOD Detection
- Using differentiable programming to obtain an energy and density-optimized exchange-correlation functional
- XDA: Accurate, Robust Disassembly with Transfer Learning
- Activated Gradients for Deep Neural Networks
- SparseDNN: Fast Sparse Deep Learning Inference on CPUs
- Multi-scale Transformer Language Models
- Generalisation of Cyberbullying Detection
- K-TanH: Efficient TanH For Deep Learning
- Improving ensemble extreme precipitation forecasts using generative artificial intelligence
- Sisyphus: A Cautionary Tale of Using Low-Degree Polynomial Activations in Privacy-Preserving Deep Learning
- TAG: Gradient Attack on Transformer-based Language Models
- "I'm Not Mad": Commonsense Implications of Negation and Contradiction
- Vision Guided Generative Pre-trained Language Models for Multimodal Abstractive Summarization
- Evaluation of self-supervised pre-training for automatic infant movement classification using wearable movement sensors
- Polyphone Disambiguation in Mandarin Chinese with Semi-Supervised Learning
- PhytNet -- Tailored Convolutional Neural Networks for Custom Botanical Data
- Salient Object Ranking with Position-Preserved Attention
- NeRV: Neural Representations for Videos
- Undivided Attention: Are Intermediate Layers Necessary for BERT?
- Pretraining the Noisy Channel Model for Task-Oriented Dialogue
- Stochastic Super-resolution of Cosmological Simulations with Denoising Diffusion Models
- A Study on MIMO Channel Estimation by 2D and 3D Convolutional Neural Networks
- Style Equalization: Unsupervised Learning of Controllable Generative Sequence Models
- Hierarchical Autoencoder-based Lossy Compression for Large-scale High-resolution Scientific Data
- Three-body renormalization group limit cycles based on unsupervised feature learning
- Demystifying BERT: Implications for Accelerator Design
- q-Neurons: Neuron Activations based on Stochastic Jackson's Derivative Operators
- End-to-End Entity Detection with Proposer and Regressor
- TLCFuse: Temporal Multi-Modality Fusion Towards Occlusion-Aware Semantic Segmentation-Aided Motion Planning
- Don't Fear Peculiar Activation Functions: EUAF and Beyond
- SMU: smooth activation function for deep networks using smoothing maximum technique
- Video Frame Interpolation Transformer
- Dual Interaction Network with Cross-Image Attention for Medical Image Segmentation
- Deep Feature Response Discriminative Calibration
- Stability and Generalization of Bilevel Programming in Hyperparameter Optimization
- PreSizE: Predicting Size in E-Commerce using Transformers
- Reborn Mechanism: Rethinking the Negative Phase Information Flow in Convolutional Neural Network
- Efficient and Accurate Gradients for Neural SDEs
- Multi-Scale Local-Temporal Similarity Fusion for Continuous Sign Language Recognition
- Sparse Attention with Linear Units
- TEASEL: A Transformer-Based Speech-Prefixed Language Model
- Initializing ReLU networks in an expressive subspace of weights
- Subquadratic Overparameterization for Shallow Neural Networks
- Unsupervised Explanation Generation for Machine Reading Comprehension
- ERNIE-M: Enhanced Multilingual Representation by Aligning Cross-lingual Semantics with Monolingual Corpora
- TransMask: A Compact and Fast Speech Separation Model Based on Transformer
- Optimized spiking neurons classify images with high accuracy through temporal coding with two spikes
- Learn-able parameter guided Activation Functions
- Scope and Arbitration in Machine Learning Clinical EEG Classification
- Orthogonal-Padé Activation Functions: Trainable Activation functions for smooth and faster convergence in deep networks
- Soft Autoencoder and Its Wavelet Adaptation Interpretation
- Octa: Omissions and Conflicts in Target-Aspect Sentiment Analysis
- Simultaneous Speech Translation for Live Subtitling: from Delay to Display
- Attention as Activation
- AMMASurv: Asymmetrical Multi-Modal Attention for Accurate Survival Analysis with Whole Slide Images and Gene Expression Data
- On the Bias Against Inductive Biases
- RaftMLP: How Much Can Be Done Without Attention and with Less Spatial Locality?
- Regularized Flexible Activation Function Combinations for Deep Neural Networks
- Universal Performance Gap of Neural Quantum States Applied to the Hofstadter-Bose-Hubbard Model
- Randomized Stochastic Gradient Descent Ascent
- Scalable neural network-based blackbox optimization
- PairConnect: A Compute-Efficient MLP Alternative to Attention
- Towards Accurate Quantization and Pruning via Data-free Knowledge Transfer
- PointMixer: MLP-Mixer for Point Cloud Understanding
- Verifying Quantized Neural Networks using SMT-Based Model Checking
- SFB-net for cardiac segmentation: Bridging the semantic gap with attention
- Structured second-order methods via natural gradient descent
- Evaluating Graphical Perception Capabilities of Vision Transformers
- SMedBERT: A Knowledge-Enhanced Pre-trained Language Model with Structured Semantics for Medical Text Mining
- Discover the Mysteries of the Maya: Selected Contributions from the Machine Learning Challenge & The Discovery Challenge Workshop at ECML PKDD 2021
- K-PLUG: Knowledge-injected Pre-trained Language Model for Natural Language Understanding and Generation in E-Commerce
- Detecting Logical Relation In Contract Clauses
- Parametric Rectified Power Sigmoid Units: Learning Nonlinear Neural Transfer Analytical Forms
- Computational homogenization for aerogel-like polydisperse open-porous materials using neural network--based surrogate models on the microscale
- Exploring the Early Universe with Deep Learning
- Using mixup as regularization and tuning hyper-parameters for ResNets
- MC-SSL0.0: Towards Multi-Concept Self-Supervised Learning
- No Answer is Better Than Wrong Answer: A Reflection Model for Document Level Machine Reading Comprehension
- Interpretable Prediction of Lymph Node Metastasis in Rectal Cancer MRI Using Variational Autoencoders
- Adversarial Generation and Encoding of Nested Texts
- Commonsense Knowledge in Word Associations and ConceptNet
- Multilingual Translation via Grafting Pre-trained Language Models
- Implicit regularization of deep residual networks towards neural ODEs
- Knowledge Transfer by Discriminative Pre-training for Academic Performance Prediction
- Towards Automatic Evaluation of Dialog Systems: A Model-Free Off-Policy Evaluation Approach
- Privacy-Preserving Public Release of Datasets for Support Vector Machine Classification
- Differentiable Random Access Memory using Lattices
- Multi-mode Transformer Transducer with Stochastic Future Context
- Sparsity-Probe: Analysis tool for Deep Learning Models
- Sines, Transient, Noise Neural Modeling of Piano Notes
- AMU-EURANOVA at CASE 2021 Task 1: Assessing the stability of multilingual BERT
- PlueckerNet: Learn to Register 3D Line Reconstructions
- ComQA:Compositional Question Answering via Hierarchical Graph Neural Networks
- CITIES: Contextual Inference of Tail-Item Embeddings for Sequential Recommendation
- EIS -- a family of activation functions combining Exponential, ISRU, and Softplus
- KST-Mixer: Kinematic Spatio-Temporal Data Mixer For Colon Shape Estimation
- Explainable Recommendations via Attentive Multi-Persona Collaborative Filtering
- ViFiT: Reconstructing Vision Trajectories from IMU and Wi-Fi Fine Time Measurements
- Introducing the DOME Activation Functions
- ISyNet: Convolutional Neural Networks design for AI accelerator
- Analysis of the Compaction Behavior of Textile Reinforcements in Low-Resolution In-Situ CT Scans via Machine-Learning and Descriptor-Based Methods
- A Lightweight Graph Transformer Network for Human Mesh Reconstruction from 2D Human Pose
- End-to-end Biomedical Entity Linking with Span-based Dictionary Matching
- MaxVA: Fast Adaptation of Step Sizes by Maximizing Observed Variance of Gradients
- jBOT: Semantic Jet Representation Clustering Emerges from Self-Distillation
- Interpretable Deep Learning for Stock Returns: A Consensus-Bottleneck Asset Pricing Model
- SocialTrans: A Deep Sequential Model with Social Information for Web-Scale Recommendation Systems
- MIRA: Leveraging Multi-Intention Co-click Information in Web-scale Document Retrieval using Deep Neural Networks
- Resolution scaling governs DINOv3 transfer performance in chest radiograph classification
- Avoiding Kernel Fixed Points: Computing with ELU and GELU Infinite Networks
- Refine Neutrino Events Reconstruction with BEiT-3