5 papers
Relaxed Attention for Transformer Models
Timo Lohrenz, Björn Möller, Zhengyang Li +1
The powerful modeling capabilities of all-attention-based transformer architectures often cause overfitting and - for natural language processing tasks - lead to an implicitly lear…
The Quadratic Wasserstein Metric With Squaring Scaling For Seismic Velocity Inversion
Zhengyang Li, Yijia Tang, Jing Chen +1
The quadratic Wasserstein metric has shown its power in measuring the difference between probability densities, which benefits optimization objective function with better convexity…
Multi-Encoder Learning and Stream Fusion for Transformer-Based End-to-End Automatic Speech Recognition
Timo Lohrenz, Zhengyang Li, Tim Fingscheidt
Stream fusion, also known as system combination, is a common technique in automatic speech recognition for traditional hybrid hidden Markov model approaches, yet mostly unexplored…
Point Spread Function Estimation for Wide Field Small Aperture Telescopes with Deep Neural Networks and Calibration Data
Peng Jia, Xuebo Wu, Zhengyang Li +4
The point spread function (PSF) reflects states of a telescope and plays an important role in development of data processing methods, such as PSF based astrometry, photometry and i…
Point Spread Function Modelling for Wide Field Small Aperture Telescopes with a Denoising Autoencoder
Peng Jia, Xiyu Li, Zhengyang Li +2
The point spread function reflects the state of an optical telescope and it is important for data post-processing methods design. For wide field small aperture telescopes, the poin…