Decoupled Weight Decay Regularization
arXiv:1711.05101
Abstract
L regularization and weight decay regularization are equivalent for standard stochastic gradient descent (when rescaled by the learning rate), but as we demonstrate this is \emph{not} the case for adaptive gradient algorithms, such as Adam. While common implementations of these algorithms employ L regularization (often calling it "weight decay" in what may be misleading due to the inequivalence we expose), we propose a simple modification to recover the original formulation of weight decay regularization by \emph{decoupling} the weight decay from the optimization steps taken w.r.t. the loss function. We provide empirical evidence that our proposed modification (i) decouples the optimal choice of weight decay factor from the setting of the learning rate for both standard SGD and Adam and (ii) substantially improves Adam's generalization performance, allowing it to compete with SGD with momentum on image classification datasets (on which it was previously typically outperformed by the latter). Our proposed decoupled weight decay has already been adopted by many researchers, and the community has implemented it in TensorFlow and PyTorch; the complete source code for our experiments is available at https://github.com/loshchil/AdamW-and-SGDW
Published as a conference paper at ICLR 2019
Cited by in corpus (1261)
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
- LoRA: Low-Rank Adaptation of Large Language Models
- Attention Mechanisms in Computer Vision: A Survey
- PVT v2: Improved Baselines with Pyramid Vision Transformer
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence
- E(3)-Equivariant Graph Neural Networks for Data-Efficient and Accurate Interatomic Potentials
- WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
- Transformer in Transformer
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
- Atomistic Line Graph Neural Network for Improved Materials Property Predictions
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
- AdaBins: Depth Estimation using Adaptive Bins
- Twins: Revisiting the Design of Spatial Attention in Vision Transformers
- Reproducible scaling laws for contrastive language-image learning
- MixMatch: A Holistic Approach to Semi-Supervised Learning
- ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision
- MedViT: A Robust Vision Transformer for Generalized Medical Image Classification
- Hyperspectral Image Classification-Traditional to Deep Models: A Survey for Future Prospects
- Conditional Positional Encodings for Vision Transformers
- Lookahead Optimizer: k steps forward, 1 step back
- TransTrack: Multiple Object Tracking with Transformer
- Early Convolutions Help Transformers See Better
- Modeling the Dynamics of PDE Systems with Physics-Constrained Deep Auto-Regressive Networks
- SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
- TERA: Self-Supervised Learning of Transformer Encoder Representation for Speech
- Compute Trends Across Three Eras of Machine Learning
- Pre-training Enhanced Spatial-temporal Graph Neural Network for Multivariate Time Series Forecasting
- MetaFormer Baselines for Vision
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
- Automated Pavement Crack Segmentation Using U-Net-based Convolutional Neural Network
- VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts
- End-to-end Temporal Action Detection with Transformer
- WILDS: A Benchmark of in-the-Wild Distribution Shifts
- Variational Diffusion Models
- Measuring Mathematical Problem Solving With the MATH Dataset
- Action Transformer: A Self-Attention Model for Short-Time Pose-Based Human Action Recognition
- MonoViT: Self-Supervised Monocular Depth Estimation with a Vision Transformer
- A Benchmark Study of Machine Learning Models for Online Fake News Detection
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- Time-to-Event Prediction with Neural Networks and Cox Regression
- Multi-Modal Self-Supervised Learning for Recommendation
- Efficient Training of Audio Transformers with Patchout
- UNETR: Transformers for 3D Medical Image Segmentation
- Real-Time Apple Detection System Using Embedded Systems With Hardware Accelerators: An Edge AI Application
- On the Stability of Fine-tuning BERT: Misconceptions, Explanations, and Strong Baselines
- TransVOD: End-to-End Video Object Detection with Spatial-Temporal Transformers
- Masked Autoencoders Are Scalable Vision Learners
- K-Net: Towards Unified Image Segmentation
- H-vmunet: High-order Vision Mamba UNet for Medical Image Segmentation
- Unifying Vision-and-Language Tasks via Text Generation
- Self-supervised representation learning from 12-lead ECG data
- On Empirical Comparisons of Optimizers for Deep Learning
- UNITER: UNiversal Image-TExt Representation Learning
- FCN-Transformer Feature Fusion for Polyp Segmentation
- A Simple Baseline for Bayesian Uncertainty in Deep Learning
- Contrastive Masked Autoencoders are Stronger Vision Learners
- UltraLight VM-UNet: Parallel Vision Mamba Significantly Reduces Parameters for Skin Lesion Segmentation
- Per-Pixel Classification is Not All You Need for Semantic Segmentation
- Change Guiding Network: Incorporating Change Prior to Guide Change Detection in Remote Sensing Imagery
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutes
- A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
- Learning for Vehicle-to-Vehicle Cooperative Perception under Lossy Communication
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive Bias
- A Label Attention Model for ICD Coding from Clinical Text
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer
- ResT: An Efficient Transformer for Visual Recognition
- No More Fine-Tuning? An Experimental Evaluation of Prompt Tuning in Code Intelligence
- A Gated Cross-domain Collaborative Network for Underwater Object Detection
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
- All Tokens Matter: Token Labeling for Training Better Vision Transformers
- Measuring Coding Challenge Competence With APPS
- Toward Transformer-Based Object Detection
- RAFT: Recurrent All-Pairs Field Transforms for Optical Flow
- Collaborative Discrepancy Optimization for Reliable Image Anomaly Localization
- Masked-attention Mask Transformer for Universal Image Segmentation
- CenterCLIP: Token Clustering for Efficient Text-Video Retrieval
- Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels
- Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling
- HRFormer: High-Resolution Transformer for Dense Prediction
- Pix2seq: A Language Modeling Framework for Object Detection
- MP-SENet: A Speech Enhancement Model with Parallel Denoising of Magnitude and Phase Spectra
- CMGAN: Conformer-based Metric GAN for Speech Enhancement
- SwinFace: A Multi-task Transformer for Face Recognition, Expression Recognition, Age Estimation and Attribute Estimation
- Practical Deep Learning with Bayesian Principles
- An Efficient Lorentz Equivariant Graph Neural Network for Jet Tagging
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
- AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed Gradients
- RaSa: Relation and Sensitivity Aware Representation Learning for Text-based Person Search
- Revisiting Deep Learning Models for Tabular Data
- 3DSAM-adapter: Holistic adaptation of SAM from 2D to 3D for promptable tumor segmentation
- Self-Supervised Learning with Swin Transformers
- Uformer: A General U-Shaped Transformer for Image Restoration
- CMTR: Cross-modality Transformer for Visible-infrared Person Re-identification
- One Model to Synthesize Them All: Multi-contrast Multi-scale Transformer for Missing Data Imputation
- ChartGPT: Leveraging LLMs to Generate Charts from Abstract Natural Language
- Large Language Model Is Not a Good Few-shot Information Extractor, but a Good Reranker for Hard Samples!
- Sparse R-CNN: End-to-End Object Detection with Learnable Proposals
- CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows
- TrAISformer -- A Transformer Network with Sparse Augmented Data Representation and Cross Entropy Loss for AIS-based Vessel Trajectory Prediction
- CycleMLP: A MLP-like Architecture for Dense Prediction
- Physically Motivated Recursively Embedded Atom Neural Networks: Incorporating Local Completeness and Nonlocality
- Dynamic Prefix-Tuning for Generative Template-based Event Extraction
- Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
- AFDet: Anchor Free One Stage 3D Object Detection
- Attention-Based Transformers for Instance Segmentation of Cells in Microstructures
- DeCLUTR: Deep Contrastive Learning for Unsupervised Textual Representations
- Cough Against COVID: Evidence of COVID-19 Signature in Cough Sounds
- TransCAM: Transformer Attention-based CAM Refinement for Weakly Supervised Semantic Segmentation
- RegionViT: Regional-to-Local Attention for Vision Transformers
- Vector-quantized Image Modeling with Improved VQGAN
- Contrast to Divide: Self-Supervised Pre-Training for Learning with Noisy Labels
- VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
- Zoom Out and Observe: News Environment Perception for Fake News Detection
- Massively Parallel Universal Linear Transformations using a Wavelength-Multiplexed Diffractive Optical Network
- Tracker Meets Night: A Transformer Enhancer for UAV Tracking
- Best Practices for Scientific Research on Neural Architecture Search
- An End-to-End Earthquake Detection Method for Joint Phase Picking and Association using Deep Learning
- MEMO: Test Time Robustness via Adaptation and Augmentation
- Closing the Generalization Gap of Adaptive Gradient Methods in Training Deep Neural Networks
- Vision Transformers, a new approach for high-resolution and large-scale mapping of canopy heights
- AS-MLP: An Axial Shifted MLP Architecture for Vision
- Taming the Power of Diffusion Models for High-Quality Virtual Try-On with Appearance Flow
- Does the Magic of BERT Apply to Medical Code Assignment? A Quantitative Study
- Stochastic Gradient Methods with Layer-wise Adaptive Moments for Training of Deep Networks
- ViscNet: Neural network for predicting the fragility index and the temperature-dependency of viscosity
- AdamP: Slowing Down the Slowdown for Momentum Optimizers on Scale-invariant Weights
- How to train your MAML
- HTR-VT: Handwritten Text Recognition with Vision Transformer
- Negativity Spreads Faster: A Large-Scale Multilingual Twitter Analysis on the Role of Sentiment in Political Communication
- Video Object Segmentation with Adaptive Feature Bank and Uncertain-Region Refinement
- Hopfield Networks is All You Need
- Prompt Learning for News Recommendation
- KLUE: Korean Language Understanding Evaluation
- Benchmarking Detection Transfer Learning with Vision Transformers
- Lip-reading with Densely Connected Temporal Convolutional Networks
- Zero-shot Text Classification With Generative Language Models
- LDMVFI: Video Frame Interpolation with Latent Diffusion Models
- Referring Transformer: A One-step Approach to Multi-task Visual Grounding
- End-to-end Wind Turbine Wake Modelling with Deep Graph Representation Learning
- FireFly: A High-Throughput Hardware Accelerator for Spiking Neural Networks with Efficient DSP and Memory Optimization
- Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora
- GATE: Graph CCA for Temporal SElf-supervised Learning for Label-efficient fMRI Analysis
- GPT-GNN: Generative Pre-Training of Graph Neural Networks
- CFN-ESA: A Cross-Modal Fusion Network with Emotion-Shift Awareness for Dialogue Emotion Recognition
- Glass Segmentation with RGB-Thermal Image Pairs
- Modular machine learning-based elastoplasticity: generalization in the context of limited data
- A newcomer's guide to deep learning for inverse design in nano-photonics
- Layer-adaptive sparsity for the Magnitude-based Pruning
- The CAMELS Multifield Dataset: Learning the Universe's Fundamental Parameters with Artificial Intelligence
- Equivariant Flows: Exact Likelihood Generative Learning for Symmetric Densities
- Center-based 3D Object Detection and Tracking
- End-to-end Autonomous Driving with Semantic Depth Cloud Mapping and Multi-agent
- Image Captioning for Effective Use of Language Models in Knowledge-Based Visual Question Answering
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
- Latent Weights Do Not Exist: Rethinking Binarized Neural Network Optimization
- Entropy-based Logic Explanations of Neural Networks
- CCTrans: Simplifying and Improving Crowd Counting with Transformer
- g2tmn at Constraint@AAAI2021: Exploiting CT-BERT and Ensembling Learning for COVID-19 Fake News Detection
- End-to-End Video Instance Segmentation with Transformers
- TreeLearn: A deep learning method for segmenting individual trees from ground-based LiDAR forest point clouds
- Asymmetric Loss For Multi-Label Classification
- Rethinking the Hyperparameters for Fine-tuning
- Combining 3D Image and Tabular Data via the Dynamic Affine Feature Map Transform
- CNN-LSTM and Transfer Learning Models for Malware Classification based on Opcodes and API Calls
- Transformers in Self-Supervised Monocular Depth Estimation with Unknown Camera Intrinsics
- TabularNet: A Neural Network Architecture for Understanding Semantic Structures of Tabular Data
- Do We Need Zero Training Loss After Achieving Zero Training Error?
- Self-Supervised Monocular Depth Estimation with Self-Reference Distillation and Disparity Offset Refinement
- Training Strategies for Improved Lip-reading
- CMT: Convolutional Neural Networks Meet Vision Transformers
- LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding
- Video Instance Segmentation using Inter-Frame Communication Transformers
- On Transferability of Prompt Tuning for Natural Language Processing
- RDP-Net: Region Detail Preserving Network for Change Detection
- T-former: An Efficient Transformer for Image Inpainting
- InterBERT: Vision-and-Language Interaction for Multi-modal Pretraining
- Efficient Modelling of Trivializing Maps for Lattice Theory Using Normalizing Flows: A First Look at Scalability
- ISTR: End-to-End Instance Segmentation with Transformers
- How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers
- ktrain: A Low-Code Library for Augmented Machine Learning
- LiDAR-aid Inertial Poser: Large-scale Human Motion Capture by Sparse Inertial and LiDAR Sensors
- CTRAN: CNN-Transformer-based Network for Natural Language Understanding
- RobustART: Benchmarking Robustness on Architecture Design and Training Techniques
- EmoNeXt: an Adapted ConvNeXt for Facial Emotion Recognition
- HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
- Unraveling Complex Data Diversity in Underwater Acoustic Target Recognition through Convolution-based Mixture of Experts
- Towards duration robust weakly supervised sound event detection
- Object DGCNN: 3D Object Detection using Dynamic Graphs
- Sparse Compressed Spiking Neural Network Accelerator for Object Detection
- Multispectral Vineyard Segmentation: A Deep Learning approach
- Luna: Linear Unified Nested Attention
- SparsePoser: Real-time Full-body Motion Reconstruction from Sparse Data
- Fully Sparse Fusion for 3D Object Detection
- Three Mechanisms of Weight Decay Regularization
- A neural operator-based surrogate solver for free-form electromagnetic inverse design
- Less is More: ClipBERT for Video-and-Language Learning via Sparse Sampling
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point Modeling
- Primordial non-Gaussianity from the Completed SDSS-IV extended Baryon Oscillation Spectroscopic Survey I: Catalogue Preparation and Systematic Mitigation
- HANNA: Hard-constraint Neural Network for Consistent Activity Coefficient Prediction
- Self-Supervised RF Signal Representation Learning for NextG Signal Classification with Deep Learning
- SwinTrack: A Simple and Strong Baseline for Transformer Tracking
- SparseSwin: Swin Transformer with Sparse Transformer Block
- EmoBERTa: Speaker-Aware Emotion Recognition in Conversation with RoBERTa
- On the Effects of Different Types of Label Noise in Multi-Label Remote Sensing Image Classification
- Coordinated Sum-Rate Maximization in Multicell MU-MIMO with Deep Unrolling
- Differentially Private Fine-tuning of Language Models
- CenterNet3D: An Anchor Free Object Detector for Point Cloud
- ViDT: An Efficient and Effective Fully Transformer-based Object Detector
- Force-Field-Enhanced Neural Network Interactions: from Local Equivariant Embedding to Atom-in-Molecule properties and long-range effects
- Reweighted Proximal Pruning for Large-Scale Language Representation
- A 23 W Keyword Spotting IC with Ring-Oscillator-Based Time-Domain Feature Extraction
- Multiway Non-rigid Point Cloud Registration via Learned Functional Map Synchronization
- Using convolutional neural networks to predict galaxy metallicity from three-color images
- A Survey of Optimization Methods from a Machine Learning Perspective
- Towards an astronomical foundation model for stars with a Transformer-based model
- Is Heterophily A Real Nightmare For Graph Neural Networks To Do Node Classification?
- Jasper: An End-to-End Convolutional Neural Acoustic Model
- Permutationless Many-Jet Event Reconstruction with Symmetry Preserving Attention Networks
- Toward Accurate Interpretable Predictions of Materials Properties within Transformer Language Models
- Mask and Reason: Pre-Training Knowledge Graph Transformers for Complex Logical Queries
- An Attention Free Transformer
- ObjFormer: Learning Land-Cover Changes From Paired OSM Data and Optical High-Resolution Imagery via Object-Guided Transformer
- Vision Transformers with Patch Diversification
- A Heterogeneous Graph-Based Multi-Task Learning for Fault Event Diagnosis in Smart Grid
- Underwater Acoustic Target Recognition based on Smoothness-inducing Regularization and Spectrogram-based Data Augmentation
- Radial Basis Function Networks for Convolutional Neural Networks to Learn Similarity Distance Metric and Improve Interpretability
- Video Moment Retrieval from Text Queries via Single Frame Annotation
- Machine learning for molecular dynamics with strongly correlated electrons
- Distilling Knowledge from Reader to Retriever for Question Answering
- Vector Quantized Diffusion Model for Text-to-Image Synthesis
- QueryProp: Object Query Propagation for High-Performance Video Object Detection
- P2C: Self-Supervised Point Cloud Completion from Single Partial Clouds
- A Recurrent Vision-and-Language BERT for Navigation
- Graph-based Modeling of Online Communities for Fake News Detection
- VALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation
- DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting
- Audio ALBERT: A Lite BERT for Self-supervised Learning of Audio Representation
- Fluid Simulation on Neural Flow Maps
- A Dataset for Answering Time-Sensitive Questions
- An Overview of Indian Spoken Language Recognition from Machine Learning Perspective
- SineKAN: Kolmogorov-Arnold Networks Using Sinusoidal Activation Functions
- 12-in-1: Multi-Task Vision and Language Representation Learning
- Masked Autoencoders for Microscopy are Scalable Learners of Cellular Biology
- Parameter Efficient Multimodal Transformers for Video Representation Learning
- Faster On-Device Training Using New Federated Momentum Algorithm
- The Lottery Ticket Hypothesis for Pre-trained BERT Networks
- Structured Prediction as Translation between Augmented Natural Languages
- Molecule Identification with Rotational Spectroscopy and Probabilistic Deep Learning
- Automated HER2 Scoring in Breast Cancer Images Using Deep Learning and Pyramid Sampling
- SpectralFormer: Rethinking Hyperspectral Image Classification with Transformers
- A cascade network for Detecting COVID-19 using chest x-rays
- MatSynth: A Modern PBR Materials Dataset
- Greedy-layer Pruning: Speeding up Transformer Models for Natural Language Processing
- Gibbs-Helmholtz Graph Neural Network: capturing the temperature dependency of activity coefficients at infinite dilution
- Mining the Benefits of Two-stage and One-stage HOI Detection
- BiToD: A Bilingual Multi-Domain Dataset For Task-Oriented Dialogue Modeling
- Descending through a Crowded Valley - Benchmarking Deep Learning Optimizers
- Finding universal relations in subhalo properties with artificial intelligence
- Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis
- Detecting Replay Attacks Using Multi-Channel Audio: A Neural Network-Based Method
- MINN: Learning the dynamics of differential-algebraic equations and application to battery modeling
- Enhancing Network Initialization for Medical AI Models Using Large-Scale, Unlabeled Natural Images
- PSLT: A Light-weight Vision Transformer with Ladder Self-Attention and Progressive Shift
- Concurrent ischemic lesion age estimation and segmentation of CT brain using a Transformer-based network
- InversionNet3D: Efficient and Scalable Learning for 3D Full Waveform Inversion
- UP-DETR: Unsupervised Pre-training for Object Detection with Transformers
- Fully Transformer Networks for Semantic Image Segmentation
- Asking and Answering Questions to Evaluate the Factual Consistency of Summaries
- Local primordial non-Gaussianity from the large-scale clustering of photometric DESI luminous red galaxies
- PeCo: Perceptual Codebook for BERT Pre-training of Vision Transformers
- MatFuse: Controllable Material Generation with Diffusion Models
- Alternating Recurrent Dialog Model with Large-scale Pre-trained Language Models
- IQAGPT: Image Quality Assessment with Vision-language and ChatGPT Models
- BinaryDuo: Reducing Gradient Mismatch in Binary Activation Network by Coupling Binary Activations
- Few-Shot Segmentation via Cycle-Consistent Transformer
- Investigating the Limitations of Transformers with Simple Arithmetic Tasks
- PointNetKL: Deep Inference for GICP Covariance Estimation in Bathymetric SLAM
- TAda! Temporally-Adaptive Convolutions for Video Understanding
- Symplectic Learning for Hamiltonian Neural Networks
- HateCheck: Functional Tests for Hate Speech Detection Models
- MetaFormer Is Actually What You Need for Vision
- ChipNet: Budget-Aware Pruning with Heaviside Continuous Approximations
- How Should Pre-Trained Language Models Be Fine-Tuned Towards Adversarial Robustness?
- Spatiotemporal Transformer for Video-based Person Re-identification
- Heterogeneous Graph Transformer
- Automatic Personalized Impression Generation for PET Reports Using Large Language Models
- ænet-PyTorch: a GPU-supported implementation for machine learning atomic potentials training
- ElasticTrainer: Speeding Up On-Device Training with Runtime Elastic Tensor Selection
- CLOOB: Modern Hopfield Networks with InfoLOOB Outperform CLIP
- DeepSVG: A Hierarchical Generative Network for Vector Graphics Animation
- CLAMP: Prompt-based Contrastive Learning for Connecting Language and Animal Pose
- Towards constraining warm dark matter with stellar streams through neural simulation-based inference
- Synthesizing Speech from Intracranial Depth Electrodes using an Encoder-Decoder Framework
- S-MLP: Spatial-Shift MLP Architecture for Vision
- Hierarchical Vision Transformers for Cardiac Ejection Fraction Estimation
- Impedance-optical Dual-modal Cell Culture Imaging with Learning-based Information Fusion
- A Memory Efficient Baseline for Open Domain Question Answering
- MST: Masked Self-Supervised Transformer for Visual Representation
- DIAS: A Dataset and Benchmark for Intracranial Artery Segmentation in DSA sequences
- MiniVLM: A Smaller and Faster Vision-Language Model
- Conditional Motion In-betweening
- DiffDance: Cascaded Human Motion Diffusion Model for Dance Generation
- Masked Particle Modeling on Sets: Towards Self-Supervised High Energy Physics Foundation Models
- Weight-Sharing Neural Architecture Search: A Battle to Shrink the Optimization Gap
- Prot2Text: Multimodal Protein's Function Generation with GNNs and Transformers
- Pixel super-resolved virtual staining of label-free tissue using diffusion models
- A Closer Look at the Robustness of Vision-and-Language Pre-trained Models
- Deep Indexed Active Learning for Matching Heterogeneous Entity Representations
- Exploring Prompt-based Few-shot Learning for Grounded Dialog Generation
- Improving Word Translation via Two-Stage Contrastive Learning
- Convolution, aggregation and attention based deep neural networks for accelerating simulations in mechanics
- StepNet: Spatial-temporal Part-aware Network for Isolated Sign Language Recognition
- N-best T5: Robust ASR Error Correction using Multiple Input Hypotheses and Constrained Decoding Space
- Denoise Pretraining on Nonequilibrium Molecules for Accurate and Transferable Neural Potentials
- Tracking perovskite crystallization via deep learning-based feature detection on 2D X-ray scattering data
- Complex-valued universal linear transformations and image encryption using spatially incoherent diffractive networks
- Training Deep Spiking Neural Networks
- CgT-GAN: CLIP-guided Text GAN for Image Captioning
- Hybrid Learning of Time-Series Inverse Dynamics Models for Locally Isotropic Robot Motion
- Enhancing MRI-Based Classification of Alzheimer's Disease with Explainable 3D Hybrid Compact Convolutional Transformers
- Delving Globally into Texture and Structure for Image Inpainting
- Transformer-Unet: Raw Image Processing with Unet
- Enhancing Heterogeneous Knowledge Graph Completion with a Novel GAT-based Approach
- DeepQMC: an open-source software suite for variational optimization of deep-learning molecular wave functions
- ATST: Audio Representation Learning with Teacher-Student Transformer
- IGLUE: A Benchmark for Transfer Learning across Modalities, Tasks, and Languages
- Accurate Deep Learning-aided Density-free Strategy for Many-Body Dispersion-corrected Density Functional Theory
- New Benchmarks for Learning on Non-Homophilous Graphs
- SynthEnsemble: A Fusion of CNN, Vision Transformer, and Hybrid Models for Multi-Label Chest X-Ray Classification
- An Energy-Efficient Spiking Neural Network for Finger Velocity Decoding for Implantable Brain-Machine Interface
- Fake or Genuine? Contextualised Text Representation for Fake Review Detection
- Well-calibrated Model Uncertainty with Temperature Scaling for Dropout Variational Inference
- X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal Transformers
- Efficient Neural Query Auto Completion
- SpellGCN: Incorporating Phonological and Visual Similarities into Language Models for Chinese Spelling Check
- AdaX: Adaptive Gradient Descent with Exponential Long Term Memory
- Sensitive Data Detection and Classification in Spanish Clinical Text: Experiments with BERT
- ProGraML: Graph-based Deep Learning for Program Optimization and Analysis
- Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep Learning
- A community-powered search of machine learning strategy space to find NMR property prediction models
- Meta-DETR: Image-Level Few-Shot Object Detection with Inter-Class Correlation Exploitation
- Unsupervised Learning of Full-Waveform Inversion: Connecting CNN and Partial Differential Equation in a Loop
- Token-Level Supervised Contrastive Learning for Punctuation Restoration
- UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding
- LTC-SUM: Lightweight Client-driven Personalized Video Summarization Framework Using 2D CNN
- Perfect is the enemy of test oracle
- Towards Robust Learning-Based Pose Estimation of Noncooperative Spacecraft
- NOPE-SAC: Neural One-Plane RANSAC for Sparse-View Planar 3D Reconstruction
- Physics-Guided Neural Networks for Intraventricular Vector Flow Mapping
- VOLO: Vision Outlooker for Visual Recognition
- Giving Commands to a Self-Driving Car: How to Deal with Uncertain Situations?
- A Comprehensive Review of State-of-The-Art Methods for Java Code Generation from Natural Language Text
- On the adequacy of untuned warmup for adaptive optimization
- Human-Adversarial Visual Question Answering
- An Empirical Study of Training End-to-End Vision-and-Language Transformers
- Object as Hotspots: An Anchor-Free 3D Object Detection Approach via Firing of Hotspots
- Navigating by Touch: Haptic Monte Carlo Localization via Geometric Sensing and Terrain Classification
- QVHighlights: Detecting Moments and Highlights in Videos via Natural Language Queries
- Meta-Learning Deep Energy-Based Memory Models
- Anti-Spoofing Using Transfer Learning with Variational Information Bottleneck
- Self-Supervised and Invariant Representations for Wireless Localization
- Constrained Neural Ordinary Differential Equations with Stability Guarantees
- Characterizing signal propagation to close the performance gap in unnormalized ResNets
- CODEBench: A Neural Architecture and Hardware Accelerator Co-Design Framework
- AngularGrad: A New Optimization Technique for Angular Convergence of Convolutional Neural Networks
- Towards Efficient Post-training Quantization of Pre-trained Language Models
- CLIP-ReIdent: Contrastive Training for Player Re-Identification
- Enquire One's Parent and Child Before Decision: Fully Exploit Hierarchical Structure for Self-Supervised Taxonomy Expansion
- Magnetohydrodynamics with Physics Informed Neural Operators
- MVT: Multi-view Vision Transformer for 3D Object Recognition
- Time Matters in Regularizing Deep Networks: Weight Decay and Data Augmentation Affect Early Learning Dynamics, Matter Little Near Convergence
- The NTT DCASE2020 Challenge Task 6 system: Automated Audio Captioning with Keywords and Sentence Length Estimation
- Deep Learning for Segmentation of Cracks in High-Resolution Images of Steel Bridges
- S3: A Spectral-Spatial Structure Loss for Pan-Sharpening Networks
- KG-BART: Knowledge Graph-Augmented BART for Generative Commonsense Reasoning
- An Explanation of In-context Learning as Implicit Bayesian Inference
- Amplifying Pathological Detection in EEG Signaling Pathways through Cross-Dataset Transfer Learning
- The Loss Surfaces of Neural Networks with General Activation Functions
- Discriminative Nearest Neighbor Few-Shot Intent Detection by Transferring Natural Language Inference
- Distribution-aware Margin Calibration for Semantic Segmentation in Images
- HOTR: End-to-End Human-Object Interaction Detection with Transformers
- Differentiable Model Compression via Pseudo Quantization Noise
- Hardware-Robust In-RRAM-Computing for Object Detection
- SafeGen: Mitigating Sexually Explicit Content Generation in Text-to-Image Models
- Landslide Detection in Real-Time Social Media Image Streams
- Global Filter Networks for Image Classification
- Associating Objects with Transformers for Video Object Segmentation
- 1st Place Solution for Waymo Open Dataset Challenge -- 3D Detection and Domain Adaptation
- Taming Visually Guided Sound Generation
- DNN-MG: A Hybrid Neural Network/Finite Element Method with Applications to 3D Simulations of the Navier-Stokes Equations
- Enhancing Few-shot Image Classification with Cosine Transformer
- Multifield Cosmology with Artificial Intelligence
- An Empirical Study of Tokenization Strategies for Various Korean NLP Tasks
- SFace: An Efficient Network for Face Detection in Large Scale Variations
- SGTBN: Generating Dense Depth Maps from Single-Line LiDAR
- Improving Pre-trained Language Model Fine-tuning with Noise Stability Regularization
- Simplifying Hamiltonian and Lagrangian Neural Networks via Explicit Constraints
- Multimap targeted free energy estimation
- ReCycle: Resilient Training of Large DNNs using Pipeline Adaptation
- BodyFormer: Semantics-guided 3D Body Gesture Synthesis with Transformer
- Cross-Domain Aspect Extraction using Transformers Augmented with Knowledge Graphs
- MUCM-Net: A Mamba Powered UCM-Net for Skin Lesion Segmentation
- LOANet: A Lightweight Network Using Object Attention for Extracting Buildings and Roads from UAV Aerial Remote Sensing Images
- Transfer Learning with Foundational Models for Time Series Forecasting using Low-Rank Adaptations
- Knowledge Amalgamation for Object Detection with Transformers
- HODOR: High-level Object Descriptors for Object Re-segmentation in Video Learned from Static Images
- RAFT-Stereo: Multilevel Recurrent Field Transforms for Stereo Matching
- Node-weighted Graph Convolutional Network for Depression Detection in Transcribed Clinical Interviews
- Connecting optical morphology, environment, and HI mass fraction for low-redshift galaxies using deep learning
- Face Morphing Attack Detection with Denoising Diffusion Probabilistic Models
- M6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter Pretraining
- EarlyBERT: Efficient BERT Training via Early-bird Lottery Tickets
- Kernelized information bottleneck leads to biologically plausible 3-factor Hebbian learning in deep networks
- Attention-aware non-rigid image registration for accelerated MR imaging
- Hypersolvers: Toward Fast Continuous-Depth Models
- Investigating Shift-Variance of Convolutional Neural Networks in Ultrasound Image Segmentation
- Learning from Context or Names? An Empirical Study on Neural Relation Extraction
- Robust Lottery Tickets for Pre-trained Language Models
- Estimating Redundancy in Clinical Text
- Multi-queue Momentum Contrast for Microvideo-Product Retrieval
- Language Modelling with Pixels
- Automated Concatenation of Embeddings for Structured Prediction
- UnifiedGesture: A Unified Gesture Synthesis Model for Multiple Skeletons
- QPIC: Query-Based Pairwise Human-Object Interaction Detection with Image-Wide Contextual Information
- On the Optimal Weighted Regularization in Overparameterized Linear Regression
- Adversarial multi-task underwater acoustic target recognition: towards robustness against various influential factors
- Regularization Matters in Policy Optimization
- From system models to class models: An in-context learning paradigm
- MORE: A Multimodal Object-Entity Relation Extraction Dataset with a Benchmark Evaluation
- CondLaneNet: a Top-to-down Lane Detection Framework Based on Conditional Convolution
- A case for new neural network smoothness constraints
- A Cascade Transformer-based Model for 3D Dose Distribution Prediction in Head and Neck Cancer Radiotherapy
- Attention Aided CSI Wireless Localization
- Unfolding AIS transmission behavior for vessel movement modeling on noisy data leveraging machine learning
- Astroconformer: The Prospects of Analyzing Stellar Light Curves with Transformer-Based Deep Learning Models
- Towards Human-like Perception: Learning Structural Causal Model in Heterogeneous Graph
- Evaluating histopathology transfer learning with ChampKit
- Deep neural networks for choice analysis: Enhancing behavioral regularity with gradient regularization
- Cooperative data-driven modeling
- Accelerating Plasmonic Hydrogen Sensors for Inert Gas Environments by Transformer-Based Deep Learning
- Keep the Gradients Flowing: Using Gradient Flow to Study Sparse Network Optimization
- Diverse Image Inpainting with Bidirectional and Autoregressive Transformers
- Multi-View Reasoning: Consistent Contrastive Learning for Math Word Problem
- UniNeXt: Exploring A Unified Architecture for Vision Recognition
- Multimodal machine learning with large language embedding model for polymer property prediction
- Variational principle to regularize machine-learned density functionals: the non-interacting kinetic-energy functional
- Eliminating polarization leakage effect for neutral hydrogen intensity mapping with deep learning
- USCORE: An Effective Approach to Fully Unsupervised Evaluation Metrics for Machine Translation
- OnDev-LCT: On-Device Lightweight Convolutional Transformers towards federated learning
- AdaViT: Adaptive Vision Transformers for Efficient Image Recognition
- IntFormer: Predicting pedestrian intention with the aid of the Transformer architecture
- NFAD: Fixing anomaly detection using normalizing flows
- Jack and Masters of all Trades: One-Pass Learning Sets of Model Sets From Large Pre-Trained Models
- Inverse molecular design and parameter optimization with Hückel theory using automatic differentiation
- Ultra Fast Speech Separation Model with Teacher Student Learning
- RAGO: Recurrent Graph Optimizer For Multiple Rotation Averaging
- Regularization Matters: A Nonparametric Perspective on Overparametrized Neural Network
- Do Language Models Have Beliefs? Methods for Detecting, Updating, and Visualizing Model Beliefs
- Vehicle Occurrence-based Parking Space Detection
- Distill-SODA: Distilling Self-Supervised Vision Transformer for Source-Free Open-Set Domain Adaptation in Computational Pathology
- DocScanner: Robust Document Image Rectification with Progressive Learning
- FLERT: Document-Level Features for Named Entity Recognition
- An Ensemble of Knowledge Sharing Models for Dynamic Hand Gesture Recognition
- Condition-Invariant Semantic Segmentation
- Improved Difference Images for Change Detection Classifiers in SAR Imagery Using Deep Learning
- TransVOS: Video Object Segmentation with Transformers
- Hate-Alert@DravidianLangTech-EACL2021: Ensembling strategies for Transformer-based Offensive language Detection
- Net: Accurate Panorama Depth Estimation on Spherical Surface
- From Universal Language Model to Downstream Task: Improving RoBERTa-Based Vietnamese Hate Speech Detection
- Search-Engine-augmented Dialogue Response Generation with Cheaply Supervised Query Production
- Objects are Different: Flexible Monocular 3D Object Detection
- Using a thousand optimization tasks to learn hyperparameter search strategies
- A Neural Span-Based Continual Named Entity Recognition Model
- LaProp: Separating Momentum and Adaptivity in Adam
- Neural Lyapunov Differentiable Predictive Control
- Automatic Differentiation-based Full Waveform Inversion with Flexible Workflows
- SpecTNT: a Time-Frequency Transformer for Music Audio
- Text2Event: Controllable Sequence-to-Structure Generation for End-to-end Event Extraction
- Multi-Scale and Multi-Layer Contrastive Learning for Domain Generalization
- Convolutional L2LFlows: Generating Accurate Showers in Highly Granular Calorimeters Using Convolutional Normalizing Flows
- Optimizer Benchmarking Needs to Account for Hyperparameter Tuning
- EfficientPhys: Enabling Simple, Fast and Accurate Camera-Based Vitals Measurement
- Encoding protein dynamic information in graph representation for functional residue identification
- MultiMAE-DER: Multimodal Masked Autoencoder for Dynamic Emotion Recognition
- Normalized Flat Minima: Exploring Scale Invariant Definition of Flat Minima for Neural Networks using PAC-Bayesian Analysis
- M6-T: Exploring Sparse Expert Models and Beyond
- A Transformer-based Neural Language Model that Synthesizes Brain Activation Maps from Free-Form Text Queries
- Audio Event-Relational Graph Representation Learning for Acoustic Scene Classification
- RG-Flow: A hierarchical and explainable flow model based on renormalization group and sparse prior
- Event Transformer+. A multi-purpose solution for efficient event data processing
- Learning To Retrieve: How to Train a Dense Retrieval Model Effectively and Efficiently
- Are we Forgetting about Compositional Optimisers in Bayesian Optimisation?
- ERICA: Improving Entity and Relation Understanding for Pre-trained Language Models via Contrastive Learning
- Generalizing MLPs With Dropouts, Batch Normalization, and Skip Connections
- Neural Ordinary Differential Equations for Model Order Reduction of Stiff Systems
- Sam2Rad: A Segmentation Model for Medical Images with Learnable Prompts
- DeViT: Deformed Vision Transformers in Video Inpainting
- Ask "Who", Not "What": Bitcoin Volatility Forecasting with Twitter Data
- MultiCapCLIP: Auto-Encoding Prompts for Zero-Shot Multilingual Visual Captioning
- NLNDE at SemEval-2023 Task 12: Adaptive Pretraining and Source Language Selection for Low-Resource Multilingual Sentiment Analysis
- Response Generation with Context-Aware Prompt Learning
- SimPLE: Similar Pseudo Label Exploitation for Semi-Supervised Classification
- SegDiff: Image Segmentation with Diffusion Probabilistic Models
- Sound Event Detection Transformer: An Event-based End-to-End Model for Sound Event Detection
- ChartSumm: A Comprehensive Benchmark for Automatic Chart Summarization of Long and Short Summaries
- A Probabilistic Autoencoder for Type Ia Supernovae Spectral Time Series
- Unified Multi-Criteria Chinese Word Segmentation with BERT
- GUing: A Mobile GUI Search Engine using a Vision-Language Model
- The Phonetic Footprint of Parkinson's Disease
- Multi-granularity Contrastive Cross-modal Collaborative Generation for End-to-End Long-term Video Question Answering
- Learning Bounds for Risk-sensitive Learning
- Exploring the Landscape of Natural Language Processing Research
- Supervised Learning on Relational Databases with Graph Neural Networks
- Double Graph Based Reasoning for Document-level Relation Extraction
- BigBIO: A Framework for Data-Centric Biomedical Natural Language Processing
- Regress Before Construct: Regress Autoencoder for Point Cloud Self-supervised Learning
- PECI-Net: Bolus segmentation from video fluoroscopic swallowing study images using preprocessing ensemble and cascaded inference
- GPTSee: Enhancing Moment Retrieval and Highlight Detection via Description-Based Similarity Features
- Guiding the underwater acoustic target recognition with interpretable contrastive learning
- Tucano: Advancing Neural Text Generation for Portuguese
- Incorporating Anatomical Awareness for Enhanced Generalizability and Progression Prediction in Deep Learning-Based Radiographic Sacroiliitis Detection
- Table Detection for Visually Rich Document Images
- Visual Transformers for Primates Classification and Covid Detection
- InforMask: Unsupervised Informative Masking for Language Model Pretraining
- Histopathological Image Classification based on Self-Supervised Vision Transformer and Weak Labels
- MM-ALT: A Multimodal Automatic Lyric Transcription System
- ClipCap: CLIP Prefix for Image Captioning
- Geometric and Physical Quantities Improve E(3) Equivariant Message Passing
- The Causal-Neural Connection: Expressiveness, Learnability, and Inference
- Muesli: Combining Improvements in Policy Optimization
- Multitask Recalibrated Aggregation Network for Medical Code Prediction
- Word Alignment by Fine-tuning Embeddings on Parallel Corpora
- INSPIRED: Toward Sociable Recommendation Dialog Systems
- Bag of Tricks for Adversarial Training
- P-CRITICAL: A Reservoir Autoregulation Plasticity Rule for Neuromorphic Hardware
- Deep Forest: Neural Network reconstruction of the Lyman-alpha forest
- HybridQA: A Dataset of Multi-Hop Question Answering over Tabular and Textual Data
- Single-/Multi-Source Cross-Lingual NER via Teacher-Student Learning on Unlabeled Data in Target Language
- The gap between theory and practice in function approximation with deep neural networks
- A Deep Learning Based Attack for The Chaos-based Image Encryption
- Beyond the Eye: A Relational Model for Early Dementia Detection Using Retinal OCTA Images
- Learning to Adapt Domain Shifts of Moral Values via Instance Weighting
- Proxy-based Item Representation for Attribute and Context-aware Recommendation
- SLOctolyzer: Fully automatic analysis toolkit for segmentation and feature extracting in scanning laser ophthalmoscopy images
- End-to-End Zero-Shot Voice Conversion with Location-Variable Convolutions
- User Training with Error Augmentation for Electromyogram-based Gesture Classification
- Continuous Speech Separation with Conformer
- Robust Unsupervised Video Anomaly Detection by Multi-Path Frame Prediction
- Text-only domain adaptation for end-to-end ASR using integrated text-to-mel-spectrogram generator
- Reformulating HOI Detection as Adaptive Set Prediction
- Diverse Title Generation for Stack Overflow Posts with Multiple Sampling Enhanced Transformer
- Dense Voxel 3D Reconstruction Using a Monocular Event Camera
- DirectMultiStep: Direct Route Generation for Multistep Retrosynthesis
- A Hierarchical Transformer with Speaker Modeling for Emotion Recognition in Conversation
- Morphosyntactic probing of multilingual BERT models
- Automated Olfactory Bulb Segmentation on High Resolutional T2-Weighted MRI
- MOT-DETR: 3D Single Shot Detection and Tracking with Transformers to build 3D representations for Agro-Food Robots
- Heterogeneous Graph Neural Networks with Post-hoc Explanations for Multi-modal and Explainable Land Use Inference
- Aspect and Opinion Term Extraction for Hotel Reviews using Transfer Learning and Auxiliary Labels
- Learning Active Subspaces and Discovering Important Features with Gaussian Radial Basis Functions Neural Networks
- GRAPPA -- A Hybrid Graph Neural Network for Predicting Pure Component Vapor Pressures
- UFO-ViT: High Performance Linear Vision Transformer without Softmax
- Automotive Object Detection via Learning Sparse Events by Spiking Neurons
- MXR-U-Nets for Real Time Hyperspectral Reconstruction
- Subhalo effective density slope measurements from HST strong lensing data with neural likelihood-ratio estimation
- Continual BERT: Continual Learning for Adaptive Extractive Summarization of COVID-19 Literature
- Unified Question Generation with Continual Lifelong Learning
- MICDIR: Multi-scale Inverse-consistent Deformable Image Registration using UNetMSS with Self-Constructing Graph Latent
- Reconciling Modern Deep Learning with Traditional Optimization Analyses: The Intrinsic Learning Rate
- Towards Foundation Models for Materials Science: The Open MatSci ML Toolkit
- Simple data balancing achieves competitive worst-group-accuracy
- "You might think about slightly revising the title": identifying hedges in peer-tutoring interactions
- CCS Explorer: Relevance Prediction, Extractive Summarization, and Named Entity Recognition from Clinical Cohort Studies
- Deep neural networks approach to microbial colony detection -- a comparative analysis
- TernaryBERT: Distillation-aware Ultra-low Bit BERT
- PTT5: Pretraining and validating the T5 model on Brazilian Portuguese data
- Enhancing Dual-Encoders with Question and Answer Cross-Embeddings for Answer Retrieval
- Goal-Oriented Multi-Task BERT-Based Dialogue State Tracker
- GradInit: Learning to Initialize Neural Networks for Stable and Efficient Training
- End-to-End Trainable Multi-Instance Pose Estimation with Transformers
- Scan-specific Self-supervised Bayesian Deep Non-linear Inversion for Undersampled MRI Reconstruction
- Finite-difference-informed graph network for solving steady-state incompressible flows on block-structured grids
- Recall and Learn: Fine-tuning Deep Pretrained Language Models with Less Forgetting
- Weakly Supervised Learning for Analyzing Political Campaigns on Facebook
- Generative modeling of spatio-temporal weather patterns with extreme event conditioning
- Automated crack propagation measurement on asphalt concrete specimens using an optical flow-based deep neural network
- IMG2SMI: Translating Molecular Structure Images to Simplified Molecular-input Line-entry System
- LibAUC: A Deep Learning Library for X-Risk Optimization
- Semantic Role Labeling as Dependency Parsing: Exploring Latent Tree Structures Inside Arguments
- Privacy-Preserving Models for Legal Natural Language Processing
- 3D-RETR: End-to-End Single and Multi-View 3D Reconstruction with Transformers
- LGESQL: Line Graph Enhanced Text-to-SQL Model with Mixed Local and Non-Local Relations
- Learning Rates as a Function of Batch Size: A Random Matrix Theory Approach to Neural Network Training
- MonoRUn: Monocular 3D Object Detection by Reconstruction and Uncertainty Propagation
- LANA: Towards Personalized Deep Knowledge Tracing Through Distinguishable Interactive Sequences
- Factorized Fourier Neural Operators
- CoCoSum: Contextual Code Summarization with Multi-Relational Graph Neural Network
- Improving Named Entity Recognition by External Context Retrieving and Cooperative Learning
- A Lightweight NMS-free Framework for Real-time Visual Fault Detection System of Freight Trains
- It Takes Two to Tango: Mixup for Deep Metric Learning
- Towards Biologically Plausible Convolutional Networks
- Positive-Negative Momentum: Manipulating Stochastic Gradient Noise to Improve Generalization
- Rethinking Token-Mixing MLP for MLP-based Vision Backbone
- A User-Friendly Framework for Generating Model-Preferred Prompts in Text-to-Image Synthesis
- Freeform surface topology prediction for prescribed illumination via semi-supervised learning
- Boosting multi-demographic federated learning for chest radiograph analysis using general-purpose self-supervised representations
- Towards Suicide Prevention from Bipolar Disorder with Temporal Symptom-Aware Multitask Learning
- An Adaptive Gradient Method with Energy and Momentum
- HisynSeg: Weakly-Supervised Histopathological Image Segmentation via Image-Mixing Synthesis and Consistency Regularization
- Improving the Accuracy of Analog-Based In-Memory Computing Accelerators Post-Training
- 3DVSR: 3D EPI Volume-based Approach for Angular and Spatial Light field Image Super-resolution
- Topic-Guided Sampling For Data-Efficient Multi-Domain Stance Detection
- Diffusion-Augmented Depth Prediction with Sparse Annotations
- Dynamic Frame Interpolation in Wavelet Domain
- FusionPainting: Multimodal Fusion with Adaptive Attention for 3D Object Detection
- Interactive Neural Painting
- HAD-Net: A Hierarchical Adversarial Knowledge Distillation Network for Improved Enhanced Tumour Segmentation Without Post-Contrast Images
- Progressive Motion Context Refine Network for Efficient Video Frame Interpolation
- Transform and Tell: Entity-Aware News Image Captioning
- SIRE: Separate Intra- and Inter-sentential Reasoning for Document-level Relation Extraction
- Physics-informed generative model for drug-like molecule conformers
- MaIL: A Unified Mask-Image-Language Trimodal Network for Referring Image Segmentation
- A Multimodal Approach to Device-Directed Speech Detection with Large Language Models
- Does Yoga Make You Happy? Analyzing Twitter User Happiness using Textual and Temporal Information
- Weakly Supervised Airway Orifice Segmentation in Video Bronchoscopy
- CATE: Computation-aware Neural Architecture Encoding with Transformers
- 3DTINC: Time-Equivariant Non-Contrastive Learning for Predicting Disease Progression from Longitudinal OCTs
- Revisiting BFloat16 Training
- SAPAG: A Self-Adaptive Privacy Attack From Gradients
- Modern French Poetry Generation with RoBERTa and GPT-2
- Improved OOD Generalization via Adversarial Training and Pre-training
- SceneGraphFusion: Incremental 3D Scene Graph Prediction from RGB-D Sequences
- Number of Attention Heads vs Number of Transformer-Encoders in Computer Vision
- Learning of viscosity functions in rarefied gas flows with physics-informed neural networks
- Identifying Exoplanets with Deep Learning. IV. Removing Stellar Activity Signals from Radial Velocity Measurements Using Neural Networks
- HybridPoint: Point Cloud Registration Based on Hybrid Point Sampling and Matching
- Parameterization of Cross-Token Relations with Relative Positional Encoding for Vision MLP
- Proxy Anchor Loss for Deep Metric Learning
- Improving Bilingual Lexicon Induction with Cross-Encoder Reranking
- Cross-Attention in Coupled Unmixing Nets for Unsupervised Hyperspectral Super-Resolution
- Fusing Context Into Knowledge Graph for Commonsense Question Answering
- Learning Vision Transformer with Squeeze and Excitation for Facial Expression Recognition
- Sound2Synth: Interpreting Sound via FM Synthesizer Parameters Estimation
- SALR: Sharpness-aware Learning Rate Scheduler for Improved Generalization
- Learning Multi-Stage Multi-Grained Semantic Embeddings for E-Commerce Search
- Image Retrieval on Real-life Images with Pre-trained Vision-and-Language Models
- E2E-VLP: End-to-End Vision-Language Pre-training Enhanced by Visual Learning
- SST: Self-training with Self-adaptive Thresholding for Semi-supervised Learning
- Learning the Geodesic Embedding with Graph Neural Networks
- DeeperForensics Challenge 2020 on Real-World Face Forgery Detection: Methods and Results
- Robust and Decomposable Average Precision for Image Retrieval
- A Comparison-Relationship-Surrogate Evolutionary Algorithm for Multi-Objective Optimization
- SkillSpan: Hard and Soft Skill Extraction from English Job Postings
- Neural Sentence Ordering Based on Constraint Graphs
- Detecting Narrative Elements in Informational Text
- Knowledge Injection into Dialogue Generation via Language Models
- OSDMamba: Enhancing Oil Spill Detection from Remote Sensing Images Using Selective State Space Model
- Cross-Lingual Knowledge Distillation for Answer Sentence Selection in Low-Resource Languages
- Grounded Situation Recognition with Transformers
- AuGPT: Auxiliary Tasks and Data Augmentation for End-To-End Dialogue with Pre-Trained Language Models
- Analysis of Twitter Users' Lifestyle Choices using Joint Embedding Model
- Order in the Court: Explainable AI Methods Prone to Disagreement
- Text-to-SQL in the Wild: A Naturally-Occurring Dataset Based on Stack Exchange Data
- PyBOP: A Python package for battery model optimisation and parameterisation
- It's All Around You: Range-Guided Cylindrical Network for 3D Object Detection
- Fre-GAN: Adversarial Frequency-consistent Audio Synthesis
- Unsupervised Pre-Training for 3D Leaf Instance Segmentation
- Artificial Neural Variability for Deep Learning: On Overfitting, Noise Memorization, and Catastrophic Forgetting
- Pose Recognition with Cascade Transformers
- Self-supervised Semi-supervised Learning for Data Labeling and Quality Evaluation
- SelfDoc: Self-Supervised Document Representation Learning
- TopicRefine: Joint Topic Prediction and Dialogue Response Generation for Multi-turn End-to-End Dialogue System
- Precise Length Control in Large Language Models
- DeRi-Bot: Learning to Collaboratively Manipulate Rigid Objects via Deformable Objects
- Open-World Lifelong Graph Learning
- A General Family of Stochastic Proximal Gradient Methods for Deep Learning
- Improving Coherence and Consistency in Neural Sequence Models with Dual-System, Neuro-Symbolic Reasoning
- Escaping Saddle Points Faster with Stochastic Momentum
- on the effectiveness of generative adversarial network on anomaly detection
- MissFormer: (In-)attention-based handling of missing observations for trajectory filtering and prediction
- High-Resolution Optical Flow from 1D Attention and Correlation
- FPANet: Frequency-based Video Demoireing using Frame-level Post Alignment
- Discourse-level Relation Extraction via Graph Pooling
- Structural Knowledge Distillation: Tractably Distilling Information for Structured Predictor
- Towards General Purpose Vision Systems
- Are the Multilingual Models Better? Improving Czech Sentiment with Transformers
- Advancing High-Resolution Video-Language Representation with Large-Scale Video Transcriptions
- aSAGA: Automatic Sleep Analysis with Gray Areas
- Calibrating and Improving Graph Contrastive Learning
- SideControl: Controlled Open-domain Dialogue Generation via Additive Side Networks
- CATs: Cost Aggregation Transformers for Visual Correspondence
- Tissue Concepts: supervised foundation models in computational pathology
- Weisfeiler and Lehman Go Paths: Learning Topological Features via Path Complexes
- Extending Lagrangian and Hamiltonian Neural Networks with Differentiable Contact Models
- Symmetric Regularization based BERT for Pair-wise Semantic Reasoning
- Self-Supervised Learning by Estimating Twin Class Distributions
- Hire-MLP: Vision MLP via Hierarchical Rearrangement
- Visually-Guided Sound Source Separation with Audio-Visual Predictive Coding
- Enhancing Language Representation with Constructional Information for Natural Language Understanding
- Chunked Autoregressive GAN for Conditional Waveform Synthesis
- Real-Time Anchor-Free Single-Stage 3D Detection with IoU-Awareness
- On the convergence of PINNs
- A new method for structural diagnostics with muon tomography and deep learning
- A Two-part Transformer Network for Controllable Motion Synthesis
- SpVOS: Efficient Video Object Segmentation with Triple Sparse Convolution
- The Battleship Approach to the Low Resource Entity Matching Problem
- 3D Multiphase Heterogeneous Microstructure Generation Using Conditional Latent Diffusion Models
- Visual Reasoning: from State to Transformation
- POPNASv2: An Efficient Multi-Objective Neural Architecture Search Technique
- Fusing Modalities by Multiplexed Graph Neural Networks for Outcome Prediction in Tuberculosis
- LightningDOT: Pre-training Visual-Semantic Embeddings for Real-Time Image-Text Retrieval
- End-to-End Chess Recognition
- On the Trade-off between Redundancy and Local Coherence in Summarization
- Homography-Based Loss Function for Camera Pose Regression
- Scaling Graph Neural Networks to Large Proteins
- Approximating Instance-Dependent Noise via Instance-Confidence Embedding
- Likelihood Training of Schrödinger Bridge using Forward-Backward SDEs Theory
- Cold-start Active Learning through Self-supervised Language Modeling
- Cross-Modal Conceptualization in Bottleneck Models
- Fine-tuning of Pre-trained Transformers for Hate, Offensive, and Profane Content Detection in English and Marathi
- Model Extraction and Adversarial Transferability, Your BERT is Vulnerable!
- FNet II: Spectral Classification of Quasars, Galaxies, Stars, and broad absorption line (BAL) Quasars
- A Survey of Spatio-Temporal EEG data Analysis: from Models to Applications
- Adaptive L2 Regularization in Person Re-Identification
- Reconstruction of Cardiac Cine MRI Using Motion-Guided Deformable Alignment and Multi-Resolution Fusion
- Inferring Actual Treatment Pathways from Patient Records
- Mention Memory: incorporating textual knowledge into Transformers through entity mention attention
- CiteFusion: An Ensemble Framework for Citation Intent Classification Harnessing Dual-Model Binary Couples and SHAP Analyses
- Multimedia Generative Script Learning for Task Planning
- All Attention U-NET for Semantic Segmentation of Intracranial Hemorrhages In Head CT Images
- FOCUS: Effective Embedding Initialization for Monolingual Specialization of Multilingual Models
- Energy-GNoME: A Living Database of Selected Materials for Energy Applications
- Temporal Sentence Grounding in Streaming Videos
- Sentiment Analysis of the COVID-related r/Depression Posts
- Towards Balanced Active Learning for Multimodal Classification
- Proto: A Neural Cocktail for Generating Appealing Conversations
- Offensive Language Identification in Low-resourced Code-mixed Dravidian languages using Pseudo-labeling
- S.T.A.R.-Track: Latent Motion Models for End-to-End 3D Object Tracking with Adaptive Spatio-Temporal Appearance Representations
- AeroReformer: Aerial Referring Transformer for UAV-based Referring Image Segmentation
- Conversations Are Not Flat: Modeling the Dynamic Information Flow across Dialogue Utterances
- LaRa: Latents and Rays for Multi-Camera Bird's-Eye-View Semantic Segmentation
- cosmosage: A Natural-Language Assistant for Cosmologists
- ViNTER: Image Narrative Generation with Emotion-Arc-Aware Transformer
- verBERT: Automating Brazilian Case Law Document Multi-label Categorization Using BERT
- Table-to-Text Generation with Pretrained Diffusion Models
- AdaSGD: Bridging the gap between SGD and Adam
- Query-Based Keyphrase Extraction from Long Documents
- Reduced Data-Driven Turbulence Closure for Capturing Long-Term Statistics
- A real-time anomaly detection method for robots based on a flexible and sparse latent space
- Improving Sentence-Level Relation Extraction through Curriculum Learning
- GAN Vocoder: Multi-Resolution Discriminator Is All You Need
- Are Training Resources Insufficient? Predict First Then Explain!
- An Empirical Study on Few-shot Knowledge Probing for Pretrained Language Models
- Dartmouth CS at WNUT-2020 Task 2: Informative COVID-19 Tweet Classification Using BERT
- Training Deep Neural Networks with Adaptive Momentum Inspired by the Quadratic Optimization
- Can we Estimate Truck Accident Risk from Telemetric Data using Machine Learning?
- Few-NERD: A Few-Shot Named Entity Recognition Dataset
- Embedding Transfer with Label Relaxation for Improved Metric Learning
- NeuralProphet: Explainable Forecasting at Scale
- A Bayesian Flow Network Framework for Chemistry Tasks
- A Unified Pruning Framework for Vision Transformers
- Hybrid Generative-Contrastive Representation Learning
- AdaGDA: Faster Adaptive Gradient Descent Ascent Methods for Minimax Optimization
- A Neural Network Perturbation Theory Based on the Born Series
- Fine-Tuning Pre-trained Language Model with Weak Supervision: A Contrastive-Regularized Self-Training Approach
- Persona Authentication through Generative Dialogue
- An Empirical Study on Hyperparameter Optimization for Fine-Tuning Pre-trained Language Models
- Towards Practical Lipreading with Distilled and Efficient Models
- Multi-Fact Correction in Abstractive Text Summarization
- From Disfluency Detection to Intent Detection and Slot Filling
- IDIAPers @ Causal News Corpus 2022: Efficient Causal Relation Identification Through a Prompt-based Few-shot Approach
- An Image Patch is a Wave: Phase-Aware Vision MLP
- Twitter User Representation Using Weakly Supervised Graph Embedding
- Physics-regularized neural network of the ideal-MHD solution operator in Wendelstein 7-X configurations
- Adaptive Multi-view Rule Discovery for Weakly-Supervised Compatible Products Prediction
- v2e: From Video Frames to Realistic DVS Events
- Context-Aware Classification of Legal Document Pages
- Intra-Batch Supervision for Panoptic Segmentation on High-Resolution Images
- Learning word-referent mappings and concepts from raw inputs
- Self-Supervised Pillar Motion Learning for Autonomous Driving
- BanglaBait: Semi-Supervised Adversarial Approach for Clickbait Detection on Bangla Clickbait Dataset
- Numerically Solving Parametric Families of High-Dimensional Kolmogorov Partial Differential Equations via Deep Learning
- Large-Scale Gradient-Free Deep Learning with Recursive Local Representation Alignment
- Deep Learning-Based Position Detection for Hydraulic Cylinders Using Scattering Parameters
- MISIM: A Neural Code Semantics Similarity System Using the Context-Aware Semantics Structure
- What Changes Can Large-scale Language Models Bring? Intensive Study on HyperCLOVA: Billions-scale Korean Generative Pretrained Transformers
- Boosting Salient Object Detection with Transformer-based Asymmetric Bilateral U-Net
- Continual Learning for Text Classification with Information Disentanglement Based Regularization
- Atomistic Graph Neural Networks for metals: Application to bcc iron
- ImageNet-21K Pretraining for the Masses
- Knodle: Modular Weakly Supervised Learning with PyTorch
- M-FAC: Efficient Matrix-Free Approximations of Second-Order Information
- Deep Learning for Rheumatoid Arthritis: Joint Detection and Damage Scoring in X-rays
- A Streaming End-to-End Framework For Spoken Language Understanding
- Neural-powered unit disk graph embedding: qubits connectivity for some QUBO problems
- S3Net: A Single Stream Structure for Depth Guided Image Relighting
- Raise a Child in Large Language Model: Towards Effective and Generalizable Fine-tuning
- NoSENSE: Learned unrolled cardiac MRI reconstruction without explicit sensitivity maps
- Pediatric brain tumor classification using digital histopathology and deep learning: evaluation of SOTA methods on a multi-center Swedish cohort
- A Hybrid Vision Transformer Approach for Mathematical Expression Recognition
- Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection
- VOMTC: Vision Objects for Millimeter and Terahertz Communications
- Triggering Dark Showers with Conditional Dual Auto-Encoders
- Unified Multi-modal Diagnostic Framework with Reconstruction Pre-training and Heterogeneity-combat Tuning
- Evaluating the structure of cognitive tasks with transfer learning
- ZRIGF: An Innovative Multimodal Framework for Zero-Resource Image-Grounded Dialogue Generation
- LinCE: A Centralized Benchmark for Linguistic Code-switching Evaluation
- Classifying CMB time-ordered data through deep neural networks
- Frequency Domain Transformer Networks for Video Prediction
- Improved Multimodal Fusion for Small Datasets with Auxiliary Supervision
- Adaptive Sampling Distributed Stochastic Variance Reduced Gradient for Heterogeneous Distributed Datasets
- A comprehensive study on the prediction reliability of graph neural networks for virtual screening
- Local Boosting for Weakly-Supervised Learning
- A Dynamic Reduction Network for Point Clouds
- Lite Training Strategies for Portuguese-English and English-Portuguese Translation
- Visually Grounded Continual Learning of Compositional Phrases
- Impact of Ground Truth Quality on Handwriting Recognition
- Beyond Distributional Hypothesis: Let Language Models Learn Meaning-Text Correspondence
- X-SCITLDR: Cross-Lingual Extreme Summarization of Scholarly Documents
- The Nonlinearity Coefficient - A Practical Guide to Neural Architecture Design
- OkwuGbé: End-to-End Speech Recognition for Fon and Igbo
- On the Convergence of Step Decay Step-Size for Stochastic Optimization
- AdvPicker: Effectively Leveraging Unlabeled Data via Adversarial Discriminator for Cross-Lingual NER
- Multi-Scale High-Resolution Vision Transformer for Semantic Segmentation
- The Brier Score under Administrative Censoring: Problems and Solutions
- BarcodeBERT: Transformers for Biodiversity Analysis
- ConDA: Unsupervised Domain Adaptation for LiDAR Segmentation via Regularized Domain Concatenation
- Accelerating BERT Inference for Sequence Labeling via Early-Exit
- Voint Cloud: Multi-View Point Cloud Representation for 3D Understanding
- Mesa: A Memory-saving Training Framework for Transformers
- On the Transformer Growth for Progressive BERT Training
- Illicit object detection in X-ray images using Vision Transformers
- Deep Neural Network Training with Frank-Wolfe
- Detecting and Exorcising Statistical Demons from Language Models with Anti-Models of Negative Data
- Robustness Challenges in Model Distillation and Pruning for Natural Language Understanding
- The Neural Data Router: Adaptive Control Flow in Transformers Improves Systematic Generalization
- A Simple and Robust Framework for Cross-Modality Medical Image Segmentation applied to Vision Transformers
- EfficientQA : a RoBERTa Based Phrase-Indexed Question-Answering System
- Multi-Attribute Relation Extraction (MARE) -- Simplifying the Application of Relation Extraction
- Temporal LiDAR Frame Prediction for Autonomous Driving
- MUNet: Motion Uncertainty-aware Semi-supervised Video Object Segmentation
- ChainCQG: Flow-Aware Conversational Question Generation
- Adversarially Regularized Policy Learning Guided by Trajectory Optimization
- Large Batch Simulation for Deep Reinforcement Learning
- Exploring Transformer Based Models to Identify Hate Speech and Offensive Content in English and Indo-Aryan Languages
- Scalable Deep Graph Clustering with Random-walk based Self-supervised Learning
- Dynamically Mitigating Data Discrepancy with Balanced Focal Loss for Replay Attack Detection
- WHO 2016 subtyping and automated segmentation of glioma using multi-task deep learning
- AdaTag: Multi-Attribute Value Extraction from Product Profiles with Adaptive Decoding
- On Pursuit of Designing Multi-modal Transformer for Video Grounding
- Discriminative Reasoning for Document-level Relation Extraction
- PhytNet -- Tailored Convolutional Neural Networks for Custom Botanical Data
- DragPoser: Motion Reconstruction from Variable Sparse Tracking Signals via Latent Space Optimization
- VidLanKD: Improving Language Understanding via Video-Distilled Knowledge Transfer
- CodeNeRF: Disentangled Neural Radiance Fields for Object Categories
- ViBERTgrid: A Jointly Trained Multi-Modal 2D Document Representation for Key Information Extraction from Documents
- LNPT: Label-free Network Pruning and Training
- Teaching deep neural networks to localize single molecules for super-resolution microscopy
- Explainable Health Risk Predictor with Transformer-based Medicare Claim Encoder
- S2VC: A Framework for Any-to-Any Voice Conversion with Self-Supervised Pretrained Representations
- Don't shoot butterfly with rifles: Multi-channel Continuous Speech Separation with Early Exit Transformer
- Neural network for multi-exponential sound energy decay analysis
- Multilingual Transformers for Product Matching -- Experiments and a New Benchmark in Polish
- Planning from Pixels using Inverse Dynamics Models
- Learning Camera Localization via Dense Scene Matching
- How Have We Reacted To The COVID-19 Pandemic? Analyzing Changing Indian Emotions Through The Lens of Twitter
- Revisiting Deep Generalized Canonical Correlation Analysis
- DMS: Deep Multi-Modal Sequence Sets with Hierarchical Modality Attention
- CAFENet: Class-Agnostic Few-Shot Edge Detection Network
- A Spatiotemporal Radar-Based Precipitation Model for Water Level Prediction and Flood Forecasting
- Invariance Measures for Neural Networks
- Benchmarking Commercial Intent Detection Services with Practice-Driven Evaluations
- Speeding up Deep Model Training by Sharing Weights and Then Unsharing
- Dialogue Graph Modeling for Conversational Machine Reading
- PowerTransformer: Unsupervised Controllable Revision for Biased Language Correction
- Learning Covariance-Based Multi-Scale Representation of Neuroimaging Measures for Alzheimer Classification
- Lightweight Gaze Estimation Model Via Fusion Global Information
- Low-Resource Multi-Granularity Academic Function Recognition Based on Multiple Prompt Knowledge
- CascadeBERT: Accelerating Inference of Pre-trained Language Models via Calibrated Complete Models Cascade
- Interactive Refinement of Cross-Lingual Word Embeddings
- Smooth Proxy-Anchor Loss for Noisy Metric Learning
- The Implicit Bias for Adaptive Optimization Algorithms on Homogeneous Neural Networks
- TE-YOLOF: Tiny and efficient YOLOF for blood cell detection
- Exploiting Contextual Information with Deep Neural Networks
- Fine-tuning Strategies for Domain Specific Question Answering under Low Annotation Budget Constraints
- Scaling Object Detection by Transferring Classification Weights
- NeuralPDR: Neural Differential Equations as surrogate models for Photodissociation Regions
- Neural optimization for quantum architectures: graph embedding problems with Distance Encoder Networks
- Leveraging Slot Descriptions for Zero-Shot Cross-Domain Dialogue State Tracking
- Doubly Adaptive Scaled Algorithm for Machine Learning Using Second-Order Information
- Learning Low-Level Causal Relations using a Simulated Robotic Arm
- Benchmarking Energy-Conserving Neural Networks for Learning Dynamics from Data
- Structured Sparse R-CNN for Direct Scene Graph Generation
- HodgeNet: Learning Spectral Geometry on Triangle Meshes
- Incorporating Question Answering-Based Signals into Abstractive Summarization via Salient Span Selection
- PointNu-Net: Keypoint-assisted Convolutional Neural Network for Simultaneous Multi-tissue Histology Nuclei Segmentation and Classification
- On the Sins of Image Synthesis Loss for Self-supervised Depth Estimation
- Augmented Natural Language for Generative Sequence Labeling
- Neural Text Generation with Artificial Negative Examples
- Open Question Answering over Tables and Text
- Localizing Infinity-shaped fishes: Sketch-guided object localization in the wild
- Advanced Graph-Based Deep Learning for Probabilistic Type Inference
- Lex-BERT: Enhancing BERT based NER with lexicons
- Camera Control at the Edge with Language Models for Scene Understanding
- Multi-Task Time Series Forecasting With Shared Attention
- Deep Data Flow Analysis
- Aemulus : Precision halo mass functions in wCDM cosmologies
- MAX: Masked Autoencoder for X-ray Fluorescence in Geological Investigation
- Speech-Forensics: Towards Comprehensive Synthetic Speech Dataset Establishment and Analysis
- Bridging Text and Crystal Structures: Literature-driven Contrastive Learning for Materials Science
- Benchmarking Deep Learning Methods for Irradiance Estimation from Sky Images with Applications to Video Prediction-Based Irradiance Nowcasting
- Modality-Agnostic Style Transfer for Holistic Feature Imputation
- Transformers in Unsupervised Structure-from-Motion
- Calibration and Uncertainty for multiRater Volume Assessment in multiorgan Segmentation (CURVAS) challenge results
- HEAR: Hearing Enhanced Audio Response for Video-grounded Dialogue
- ATHENA: Mathematical Reasoning with Thought Expansion
- Fokker-Planck Score Learning: Efficient Free-Energy Estimation under Periodic Boundary Conditions
- Cognate Transformer for Automated Phonological Reconstruction and Cognate Reflex Prediction
- AViTMP: A Tracking-Specific Transformer for Single-Branch Visual Tracking
- Vortex-Induced Drag Forecast for Cylinder in Non-uniform Inflow
- On Bilingual Lexicon Induction with Large Language Models
- Application-driven Validation of Posteriors in Inverse Problems
- Fabricator: An Open Source Toolkit for Generating Labeled Training Data with Teacher LLMs
- Enhancing Rotated Object Detection via Anisotropic Gaussian Bounding Box and Bhattacharyya Distance
- Improved Training for End-to-End Streaming Automatic Speech Recognition Model with Punctuation
- FROG: A new people detection dataset for knee-high 2D range finders
- The Unreasonable Effectiveness of Large Language-Vision Models for Source-free Video Domain Adaptation
- The StatCan Dialogue Dataset: Retrieving Data Tables through Conversations with Genuine Intents
- Unsupervised Improvement of Factual Knowledge in Language Models
- Investigating Strategies for Clause Recommendation
- Novel View Synthesis of Humans using Differentiable Rendering
- Towards End-to-End Open Conversational Machine Reading
- Radial Prediction Domain Adaption Classifier for the MIDOG 2022 Challenge
- A Fast Knowledge Distillation Framework for Visual Recognition
- Local Slot Attention for Vision-and-Language Navigation
- PSG: Prompt-based Sequence Generation for Acronym Extraction
- Which Discriminator for Cooperative Text Generation?
- A Multi-level Alignment Training Scheme for Video-and-Language Grounding
- Structural Pre-training for Dialogue Comprehension
- Cross-lingual Text Classification with Heterogeneous Graph Neural Network
- CoSQA: 20,000+ Web Queries for Code Search and Question Answering
- Rectangular Flows for Manifold Learning
- On Sample Based Explanation Methods for NLP:Efficiency, Faithfulness, and Semantic Evaluation
- GroupBERT: Enhanced Transformer Architecture with Efficient Grouped Structures
- Nested and Balanced Entity Recognition using Multi-Task Learning
- A Histopathology Study Comparing Contrastive Semi-Supervised and Fully Supervised Learning
- A Stronger Baseline for Ego-Centric Action Detection
- DSEE: Dually Sparsity-embedded Efficient Tuning of Pre-trained Language Models
- Scaling Law for Recommendation Models: Towards General-purpose User Representations
- An Improved Model for Voicing Silent Speech
- HySPA: Hybrid Span Generation for Scalable Text-to-Graph Extraction
- NomMer: Nominate Synergistic Context in Vision Transformer for Visual Recognition
- 1st Place Solution for YouTubeVOS Challenge 2021:Video Instance Segmentation
- Semi-weakly Supervised Contrastive Representation Learning for Retinal Fundus Images
- Pre-training Language Model Incorporating Domain-specific Heterogeneous Knowledge into A Unified Representation
- A Transformer-based Math Language Model for Handwritten Math Expression Recognition
- Towards Structured Dynamic Sparse Pre-Training of BERT
- Trans-Encoder: Unsupervised sentence-pair modelling through self- and mutual-distillations
- Generalization Error Analysis of Neural networks with Gradient Based Regularization
- Dynamic Sliding Window for Meeting Summarization
- DILBERT: Customized Pre-Training for Domain Adaptation withCategory Shift, with an Application to Aspect Extraction
- Generate & Rank: A Multi-task Framework for Math Word Problems
- Ensemble Fine-tuned mBERT for Translation Quality Estimation
- Zero-Shot Dialogue State Tracking via Cross-Task Transfer
- Long-Short Temporal Contrastive Learning of Video Transformers
- Hierarchical Representation Learning for Markov Decision Processes
- Deep Spiking Neural Networks with Resonate-and-Fire Neurons
- Initialization and Regularization of Factorized Neural Layers
- MRCBert: A Machine Reading ComprehensionApproach for Unsupervised Summarization
- GPT3Mix: Leveraging Large-scale Language Models for Text Augmentation
- On the One-sided Convergence of Adam-type Algorithms in Non-convex Non-concave Min-max Optimization
- TS-CAM: Token Semantic Coupled Attention Map for Weakly Supervised Object Localization
- Alignment Attention by Matching Key and Query Distributions
- RoMA: Robust Model Adaptation for Offline Model-based Optimization
- Multi-layer Feature Aggregation for Deep Scene Parsing Models
- Chinese Medical Question Answer Matching Based on Interactive Sentence Representation Learning
- Ranking Neural Checkpoints
- Towards a Universal Continuous Knowledge Base
- Task-Adaptive Feature Transformer for Few-Shot Segmentation
- ALPaCA vs. GP-based Prior Learning: A Comparison between two Bayesian Meta-Learning Algorithms
- Cue-word Driven Neural Response Generation with a Shrinking Vocabulary
- Tatum-Level Drum Transcription Based on a Convolutional Recurrent Neural Network with Language Model-Based Regularized Training
- Rewriting Meaningful Sentences via Conditional BERT Sampling and an application on fooling text classifiers
- Exploring Pair-Wise NMT for Indian Languages
- Neural Language Modeling for Contextualized Temporal Graph Generation
- Content Selection Network for Document-grounded Retrieval-based Chatbots
- MUSE: Multi-Scale Temporal Features Evolution for Knowledge Tracing
- Reweighting Augmented Samples by Minimizing the Maximal Expected Loss
- Auction learning as a two-player game
- Unsupervised Geometric Disentanglement for Surfaces via CFAN-VAE
- Learning from Aggregate Observations
- Improving Irregularly Sampled Time Series Learning with Dense Descriptors of Time
- On the Trend-corrected Variant of Adaptive Stochastic Optimization Methods
- Learning scale-variant features for robust iris authentication with deep learning based ensemble framework
- Finding New Diagnostic Information for Detecting Glaucoma using Neural Networks
- Time-Delay Momentum: A Regularization Perspective on the Convergence and Generalization of Stochastic Momentum for Deep Learning
- Regime Learning for Differentiable Particle Filters
- SituationalLLM: Proactive language models with scene awareness for dynamic, contextual task guidance
- Mixture-of-Partitions: Infusing Large Biomedical Knowledge Graphs into BERT
- R&R: Metric-guided Adversarial Sentence Generation
- Fine-Grained Classroom Activity Detection from Audio with Neural Networks
- sDREAMER: Self-distilled Mixture-of-Modality-Experts Transformer for Automatic Sleep Staging
- Temporal Adaptation of BERT and Performance on Downstream Document Classification: Insights from Social Media
- Data Augmentation for Cross-Domain Named Entity Recognition
- ResIST: Layer-Wise Decomposition of ResNets for Distributed Training
- Weakly-supervised Text Classification Based on Keyword Graph
- Construction Cost Index Forecasting: A Multi-feature Fusion Approach
- Dynamic Compositionality in Recursive Neural Networks with Structure-aware Tag Representations
- Understanding of Emotion Perception from Art
- A Transformer-based Autoregressive Decoder Architecture for Hierarchical Text Classification
- BERT4SO: Neural Sentence Ordering by Fine-tuning BERT
- RaftMLP: How Much Can Be Done Without Attention and with Less Spatial Locality?
- Word Sense Linking: Disambiguating Outside the Sandbox
- On the ability of monolingual models to learn language-agnostic representations
- Vitruvion: A Generative Model of Parametric CAD Sketches
- PyEuroVoc: A Tool for Multilingual Legal Document Classification with EuroVoc Descriptors
- Sequential Reptile: Inter-Task Gradient Alignment for Multilingual Learning
- Neural Mask Generator: Learning to Generate Adaptive Word Maskings for Language Model Adaptation
- Privacy Aware Person Detection in Surveillance Data
- Image-Based Parking Space Occupancy Classification: Dataset and Baseline
- Learning invariance preserving moment closure model for Boltzmann-BGK equation
- Making Document-Level Information Extraction Right for the Right Reasons
- Distribution-aware Margin Calibration for Medical Image Segmentation
- Pose2RGBD. Generating Depth and RGB images from absolute positions
- Physics-Informed Neural State Space Models via Learning and Evolution
- L2M: Practical posterior Laplace approximation with optimization-driven second moment estimation
- Instance Segmentation XXL-CT Challenge of a Historic Airplane
- AdaL: Adaptive Gradient Transformation Contributes to Convergences and Generalizations
- NeuroBack: Improving CDCL SAT Solving using Graph Neural Networks
- AMPLIFY:Attention-based Mixup for Performance Improvement and Label Smoothing in Transformer
- Using Self-Supervised Feature Extractors with Attention for Automatic COVID-19 Detection from Speech
- Communication-Compressed Adaptive Gradient Method for Distributed Nonconvex Optimization
- HNC: Leveraging Hard Negative Captions towards Models with Fine-Grained Visual-Linguistic Comprehension Capabilities
- Arch-Net: Model Distillation for Architecture Agnostic Model Deployment
- Help! Need Advice on Identifying Advice
- On the Robustness of Pretraining and Self-Supervision for a Deep Learning-based Analysis of Diabetic Retinopathy
- Designing Pre-training Datasets from Unlabeled Data for EEG Classification with Transformers
- BERTnesia: Investigating the capture and forgetting of knowledge in BERT
- C2C-GenDA: Cluster-to-Cluster Generation for Data Augmentation of Slot Filling
- ARCA23K: An audio dataset for investigating open-set label noise
- Sample Efficient Social Navigation Using Inverse Reinforcement Learning
- VL-LTR: Learning Class-wise Visual-Linguistic Representation for Long-Tailed Visual Recognition
- Multimodal Transformer with Variable-length Memory for Vision-and-Language Navigation
- LiRA: Learning Visual Speech Representations from Audio through Self-supervision
- Pre-training with Meta Learning for Chinese Word Segmentation
- Self-Distilled Self-Supervised Representation Learning
- Quickly Finding a Benign Region via Heavy Ball Momentum in Non-Convex Optimization
- Understanding the Disharmony between Weight Normalization Family and Weight Decay: shifted Regularizer
- Deep-learning in the bioimaging wild: Handling ambiguous data with deepflash2
- TD3Net: A temporal densely connected multi-dilated convolutional network for lipreading
- Learning Co-Speech Gesture for Multimodal Aphasia Type Detection
- One to Transfer All: A Universal Transfer Framework for Vision Foundation Model with Few Data
- Momentum Centering and Asynchronous Update for Adaptive Gradient Methods
- Variational Learning for Unsupervised Knowledge Grounded Dialogs
- Rethinking Noisy Label Models: Labeler-Dependent Noise with Adversarial Awareness
- Accelerated Almost-Sure Convergence Rates for Nonconvex Stochastic Gradient Descent using Stochastic Learning Rates
- An Objective Measure of Quality for Time-Scale Modification of Audio
- Cross-language Sentence Selection via Data Augmentation and Rationale Training
- Bi-Granularity Contrastive Learning for Post-Training in Few-Shot Scene
- Compositional Modeling of Nonlinear Dynamical Systems with ODE-based Random Features
- Training With Data Dependent Dynamic Learning Rates
- Revisiting Explicit Regularization in Neural Networks for Well-Calibrated Predictive Uncertainty
- MUSER: MUltimodal Stress Detection using Emotion Recognition as an Auxiliary Task
- SimCLAD: A Simple Framework for Contrastive Learning of Acronym Disambiguation
- Uppsala NLP at SemEval-2021 Task 2: Multilingual Language Models for Fine-tuning and Feature Extraction in Word-in-Context Disambiguation
- NorDial: A Preliminary Corpus of Written Norwegian Dialect Use
- Traversing the Subspace of Adversarial Patches
- G-SemTMO: Tone Mapping with a Trainable Semantic Graph
- Deep Multi-Modal Sets
- Informative Sample-Aware Proxy for Deep Metric Learning
- SpartQA: : A Textual Question Answering Benchmark for Spatial Reasoning
- UU-Tax at SemEval-2022 Task 3: Improving the generalizability of language models for taxonomy classification through data augmentation
- TorontoCL at CMCL 2021 Shared Task: RoBERTa with Multi-Stage Fine-Tuning for Eye-Tracking Prediction
- Adversarial AutoMixup
- Discover the Mysteries of the Maya: Selected Contributions from the Machine Learning Challenge & The Discovery Challenge Workshop at ECML PKDD 2021
- Varianceflow: High-Quality and Controllable Text-to-Speech using Variance Information via Normalizing Flow
- Improving BERT Pretraining with Syntactic Supervision
- GIPA: A General Information Propagation Algorithm for Graph Learning
- Hierarchical Entity Typing via Multi-level Learning to Rank
- printf: Preference Modeling Based on User Reviews with Item Images and Textual Information via Graph Learning
- Discrete Point Flow Networks for Efficient Point Cloud Generation
- LSM: Learning Subspace Minimization for Low-level Vision
- Identifying the Correlation Between Language Distance and Cross-Lingual Transfer in a Multilingual Representation Space
- NTIRE 2020 Challenge on Spectral Reconstruction from an RGB Image
- Neural Subgraph Isomorphism Counting
- Differential Privacy, Linguistic Fairness, and Training Data Influence: Impossibility and Possibility Theorems for Multilingual Language Models
- Interleaved Multitask Learning with Energy Modulated Learning Progress
- Response Generation in Longitudinal Dialogues: Which Knowledge Representation Helps?
- Generative Adversarial Networks for photo to Hayao Miyazaki style cartoons
- Knowing-how & Knowing-that: A New Task for Machine Comprehension of User Manuals
- CLC: Complex Linear Coding for the DNS 2020 Challenge
- ViTAS: Vision Transformer Architecture Search
- Image interpretation by iterative bottom-up top-down processing
- Cross-document Event Identity via Dense Annotation
- Improving Cross-Lingual Reading Comprehension with Self-Training
- SIPSA-Net: Shift-Invariant Pan Sharpening with Moving Object Alignment for Satellite Imagery
- Universal Online Convex Optimization Meets Second-order Bounds
- Polygonal Unadjusted Langevin Algorithms: Creating stable and efficient adaptive algorithms for neural networks
- FoveaTer: Foveated Transformer for Image Classification
- ConvFiT: Conversational Fine-Tuning of Pretrained Language Models
- Stylized Story Generation with Style-Guided Planning
- BERTweetFR : Domain Adaptation of Pre-Trained Language Models for French Tweets
- Learning from Multiple Noisy Partial Labelers
- A Matrix Autoencoder Framework to Align the Functional and Structural Connectivity Manifolds as Guided by Behavioral Phenotypes
- Jointly Learning to Align and Translate with Transformer Models
- Opinion Prediction with User Fingerprinting
- Analysis of the Penn Korean Universal Dependency Treebank (PKT-UD): Manual Revision to Build Robust Parsing Model in Korean
- Super-resolving Dark Matter Halos using Generative Deep Learning
- EEG-Driven Image Reconstruction with Saliency-Guided Diffusion Models
- Towards Continual Entity Learning in Language Models for Conversational Agents
- Efficient Kilometer-Scale Precipitation Downscaling with Conditional Wavelet Diffusion
- Volumization as a Natural Generalization of Weight Decay
- LMVE at SemEval-2020 Task 4: Commonsense Validation and Explanation using Pretraining Language Model
- Can images help recognize entities? A study of the role of images for Multimodal NER
- Are Compressed Language Models Less Subgroup Robust?
- Understanding How Over-Parametrization Leads to Acceleration: A case of learning a single teacher neuron
- Understanding Modern Techniques in Optimization: Frank-Wolfe, Nesterov's Momentum, and Polyak's Momentum
- A Modular Analysis of Provable Acceleration via Polyak's Momentum: Training a Wide ReLU Network and a Deep Linear Network
- A Corpus of Controlled Opinionated and Knowledgeable Movie Discussions for Training Neural Conversation Models
- Multi-illuminant Color Constancy via Multi-scale Illuminant Estimation and Fusion
- MTAdam: Automatic Balancing of Multiple Training Loss Terms
- Three Pillars improving Vision Foundation Model Distillation for Lidar
- Optimal Sampling Density for Nonparametric Regression
- CNN-based real-time 2D-3D deformable registration from a single X-ray projection
- UIUC_BioNLP at SemEval-2021 Task 11: A Cascade of Neural Models for Structuring Scholarly NLP Contributions
- ESCOXLM-R: Multilingual Taxonomy-driven Pre-training for the Job Market Domain
- Unsupervised Extractive Summarization by Human Memory Simulation
- DIVE: Towards Descriptive and Diverse Visual Commonsense Generation
- A Transformer Based Pitch Sequence Autoencoder with MIDI Augmentation
- Towards Comprehensive Monocular Depth Estimation: Multiple Heads Are Better Than One
- Learning from Multiple Sources for Data-to-Text and Text-to-Data
- WASSA@IITK at WASSA 2021: Multi-task Learning and Transformer Finetuning for Emotion Classification and Empathy Prediction
- Extracting and filtering paraphrases by bridging natural language inference and paraphrasing
- BioLeaF: A Bio-plausible Learning Framework for Training of Spiking Neural Networks
- Symbolic Music Loop Generation with VQ-VAE
- Spectral Transform Forms Scalable Transformer
- ComQA:Compositional Question Answering via Hierarchical Graph Neural Networks
- Training Multimodal Systems for Classification with Multiple Objectives
- Object Tracking by Detection with Visual and Motion Cues
- Global Structure-Aware Drum Transcription Based on Self-Attention Mechanisms
- Character Transformations for Non-Autoregressive GEC Tagging
- How semantic and geometric information mutually reinforce each other in ToF object localization
- Improving Unsupervised Question Answering via Summarization-Informed Question Generation
- Invexifying Regularization of Non-Linear Least-Squares Problems
- Variance Reduction in Deep Learning: More Momentum is All You Need
- Efficient Attribute Injection for Pretrained Language Models
- DessiLBI: Exploring Structural Sparsity of Deep Networks via Differential Inclusion Paths
- Gradient-based Hyperparameter Optimization Over Long Horizons
- Enhancing Audio Augmentation Methods with Consistency Learning
- MapReader: A Computer Vision Pipeline for the Semantic Exploration of Maps at Scale
- A Novel Information-Theoretic Objective to Disentangle Representations for Fair Classification
- Automatic travel pattern extraction from visa page stamps using CNN models
- Context Normalization Layer with Applications
- Exploiting Attention-based Sequence-to-Sequence Architectures for Sound Event Localization
- Awakening Latent Grounding from Pretrained Language Models for Semantic Parsing
- Beyond the Tip of the Iceberg: Assessing Coherence of Text Classifiers
- Open-Set Representation Learning through Combinatorial Embedding
- MIPT-NSU-UTMN at SemEval-2021 Task 5: Ensembling Learning with Pre-trained Language Models for Toxic Spans Detection
- Weakly-Supervised Monocular Depth Estimationwith Resolution-Mismatched Data
- Segmentation-Based Bounding Box Generation for Omnidirectional Pedestrian Detection
- Dynamic Knowledge Distillation for Pre-trained Language Models
- Adversarial Mixture Of Experts with Category Hierarchy Soft Constraint
- Improving Multimodal fusion via Mutual Dependency Maximisation
- Table-based Fact Verification with Salience-aware Learning
- MapRE: An Effective Semantic Mapping Approach for Low-resource Relation Extraction
- Generation of tubular and membranous shape textures with curvature functionals
- Neural Ray-Tracing: Learning Surfaces and Reflectance for Relighting and View Synthesis
- Automatically Exposing Problems with Neural Dialog Models
- To Share or not to Share: Predicting Sets of Sources for Model Transfer Learning
- MarioNette: Self-Supervised Sprite Learning
- X-SRL: A Parallel Cross-Lingual Semantic Role Labeling Dataset
- Controllable Neural Dialogue Summarization with Personal Named Entity Planning
- Understanding Emotion Valence is a Joint Deep Learning Task
- Direct reconstruction of the Reionization history from 21cm 2D Power Spectra
- Learning from Multiple Noisy Augmented Data Sets for Better Cross-Lingual Spoken Language Understanding
- Classification-based Quality Estimation: Small and Efficient Models for Real-world Applications
- Text2Brain: Synthesis of Brain Activation Maps from Free-form Text Query
- Emotion Stimulus Detection in German News Headlines
- Contrastive Domain Adaptation for Question Answering using Limited Text Corpora
- Single-dataset Experts for Multi-dataset Question Answering
- Finetuning Pretrained Transformers into Variational Autoencoders
- CAST: Enhancing Code Summarization with Hierarchical Splitting and Reconstruction of Abstract Syntax Trees
- On the Significance of Question Encoder Sequence Model in the Out-of-Distribution Performance in Visual Question Answering
- Want to Identify, Extract and Normalize Adverse Drug Reactions in Tweets? Use RoBERTa
- Towards Explainable Fact Checking
- TLDR9+: A Large Scale Resource for Extreme Summarization of Social Media Posts
- It's FLAN time! Summing feature-wise latent representations for interpretability
- Robust Generalization Strategies for Morpheme Glossing in an Endangered Language Documentation Context
- Camera Agnostic Two-Head Network for Ego-Lane Inference
- Learning to Follow Language Instructions with Compositional Policies
- Improving Multi-Party Dialogue Discourse Parsing via Domain Integration
- Comparing Classes of Estimators: When does Gradient Descent Beat Ridge Regression in Linear Models?
- GIPFA: Generating IPA Pronunciation from Audio
- Generation and Simulation of Yeast Microscopy Imagery with Deep Learning
- Modulating Regularization Frequency for Efficient Compression-Aware Model Training
- MaxVA: Fast Adaptation of Step Sizes by Maximizing Observed Variance of Gradients
- CoRI: Collective Relation Integration with Data Augmentation for Open Information Extraction
- Measurement of Hybrid Rocket Solid Fuel Regression Rate for a Slab Burner using Deep Learning
- Netmarble AI Center's WMT21 Automatic Post-Editing Shared Task Submission
- MOLUCINATE: A Generative Model for Molecules in 3D Space
- Evaluating Transferability of BERT Models on Uralic Languages
- Reducing Exposure Bias in Training Recurrent Neural Network Transducers
- Deploying a BERT-based Query-Title Relevance Classifier in a Production System: a View from the Trenches
- SciXGen: A Scientific Paper Dataset for Context-Aware Text Generation
- Exploring Task Difficulty for Few-Shot Relation Extraction
- Tiered Reasoning for Intuitive Physics: Toward Verifiable Commonsense Language Understanding
- AMU-EURANOVA at CASE 2021 Task 1: Assessing the stability of multilingual BERT
- Tensor Normal Training for Deep Learning Models
- Multi-view 3D Reconstruction with Transformer
- Modular Self-Supervision for Document-Level Relation Extraction
- An Explicit-Joint and Supervised-Contrastive Learning Framework for Few-Shot Intent Classification and Slot Filling
- Sinusoidal Flow: A Fast Invertible Autoregressive Flow
- Towards Incremental Transformers: An Empirical Analysis of Transformer Models for Incremental NLU
- Retrieval Enhanced Model for Commonsense Generation
- SQALER: Scaling Question Answering by Decoupling Multi-Hop and Logical Reasoning
- Data-to-text Generation by Splicing Together Nearest Neighbors
- Transforming Multi-Conditioned Generation from Meaning Representation
- REPT: Bridging Language Models and Machine Reading Comprehension via Retrieval-Based Pre-training
- Spatially Constrained Transformer with Efficient Global Relation Modelling for Spatio-Temporal Prediction
- MIX : a Multi-task Learning Approach to Solve Open-Domain Question Answering
- Connect-the-Dots: Bridging Semantics between Words and Definitions via Aligning Word Sense Inventories
- Incorporating Connections Beyond Knowledge Embeddings: A Plug-and-Play Module to Enhance Commonsense Reasoning in Machine Reading Comprehension
- TATL at W-NUT 2020 Task 2: A Transformer-based Baseline System for Identification of Informative COVID-19 English Tweets
- Generative Optimization Networks for Memory Efficient Data Generation
- MeLT: Message-Level Transformer with Masked Document Representations as Pre-Training for Stance Detection
- MirrorWiC: On Eliciting Word-in-Context Representations from Pretrained Language Models
- SERE: Exploring Feature Self-relation for Self-supervised Transformer
- Dialogue Inspectional Summarization with Factual Inconsistency Awareness
- P-Adapters: Robustly Extracting Factual Information from Language Models with Diverse Prompts
- Aspect-Based Argument Mining
- Attention-Free Keyword Spotting
- BERT Has Uncommon Sense: Similarity Ranking for Word Sense BERTology
- Few-Shot Transformation of Common Actions into Time and Space
- Simple and Effective Input Reformulations for Translation
- Neural network emulator to constrain the high- IGM thermal state from Lyman- forest flux auto-correlation function
- 3rd Place Solution for NeurIPS 2021 Shifts Challenge: Vehicle Motion Prediction
- Ripple Attention for Visual Perception with Sub-quadratic Complexity
- The Effectiveness of Intermediate-Task Training for Code-Switched Natural Language Understanding
- MaskBEV: Joint Object Detection and Footprint Completion for Bird's-eye View 3D Point Clouds
- ContourRend: A Segmentation Method for Improving Contours by Rendering
- A New Adaptive Gradient Method with Gradient Decomposition
- SFTrack++: A Fast Learnable Spectral Segmentation Approach for Space-Time Consistent Tracking
- Coarse-to-Fine Memory Matching for Joint Retrieval and Classification
- Noise Stability Regularization for Improving BERT Fine-tuning
- Efficient transfer learning for NLP with ELECTRA
- Semi-Relaxed Quantization with DropBits: Training Low-Bit Neural Networks via Bit-wise Regularization
- Trident Pyramid Networks: The importance of processing at the feature pyramid level for better object detection
- Event-Driven Learning of Systematic Behaviours in Stock Markets
- UniRE: A Unified Label Space for Entity Relation Extraction
- Disentangling Semantics and Syntax in Sentence Embeddings with Pre-trained Language Models
- Knowledge-driven Active Learning
- Dialogue State Tracking with a Language Model using Schema-Driven Prompting
- Vision Pair Learning: An Efficient Training Framework for Image Classification
- Exact Backpropagation in Binary Weighted Networks with Group Weight Transformations
- Adaptive Gradient Method with Resilience and Momentum
- The Second Place Solution for ICCV2021 VIPriors Instance Segmentation Challenge
- Softer Pruning, Incremental Regularization
- Speech Imagery Classification using Length-Wise Training based on Deep Learning
- Efficient fine-tuning of 37-level GraphCast with the Canadian global deterministic analysis
- Deep Learning Based Vehicle Tracking System Using License Plate Detection And Recognition