2 citations · 7 across the 5 of their papers we have counts for
8 papers
A Transformer-based Cross-modal Fusion Model with Adversarial Training for VQA Challenge 2021
Ke-Han Lu, Bo-Han Fang, Kuan-Yu Chen
In this paper, inspired by the successes of visionlanguage pre-trained models and the benefits from training with adversarial attacks, we present a novel transformerbased cross-mod…
Speech Recognition by Simply Fine-tuning BERT
Wen-Chin Huang, Chia-Hua Wu, Shang-Bao Luo +3
We propose a simple method for automatic speech recognition (ASR) by fine-tuning BERT, which is a language model (LM) trained on large-scale unlabeled text data and can generate ri…
Investigation of Sentiment Controllable Chatbot
Hung-yi Lee, Cheng-Hao Ho, Chien-Fu Lin +5
Conventional seq2seq chatbot models attempt only to find sentences with the highest probabilities conditioned on the input sequences, without considering the sentiment of the outpu…
An Audio-enriched BERT-based Framework for Spoken Multiple-choice Question Answering
Chia-Chih Kuo, Shang-Bao Luo, Kuan-Yu Chen
In a spoken multiple-choice question answering (SMCQA) task, given a passage, a question, and multiple choices all in the form of speech, the machine needs to pick the correct choi…
A neural document language modeling framework for spoken document retrieval
Li-Phen Yen, Zhen-Yu Wu, Kuan-Yu Chen
Recent developments in deep learning have led to a significant innovation in various classic and practical subjects, including speech recognition, computer vision, question answeri…
Completely Unsupervised Speech Recognition By A Generative Adversarial Network Harmonized With Iteratively Refined Hidden Markov Models
Kuan-Yu Chen, Che-Ping Tsai, Da-Rong Liu +2
Producing a large annotated speech corpus for training ASR systems remains difficult for more than 95% of languages all over the world which are low-resourced, but collecting a rel…