4 citations · 5 across the 3 of their papers we have counts for
3 papers
g2pW: A Conditional Weighted Softmax BERT for Polyphone Disambiguation in Mandarin
Yi-Chang Chen, Yu-Chuan Chang, Yen-Cheng Chang +1
Polyphone disambiguation is the most crucial task in Mandarin grapheme-to-phoneme (g2p) conversion. Previous studies have approached this problem using pre-trained language models,…
Traditional Chinese Synthetic Datasets Verified with Labeled Data for Scene Text Recognition
Yi-Chang Chen, Yu-Chuan Chang, Yen-Cheng Chang +1
Scene text recognition (STR) has been widely studied in academia and industry. Training a text recognition model often requires a large amount of labeled data, but data labeling ca…
OOWL500: Overcoming Dataset Collection Bias in the Wild
Brandon Leung, Chih-Hui Ho, Amir Persekian +5
The hypothesis that image datasets gathered online "in the wild" can produce biased object recognizers, e.g. preferring professional photography or certain viewing angles, is studi…