2 papers
cs.CV2024
Sharingan: Extract User Action Sequence from Desktop Recordings
Yanting Chen, Yi Ren, Xiaoting Qin +7
Video recordings of user activities, particularly desktop recordings, offer a rich source of data for understanding user behaviors and automating processes. However, despite advanc…
cs.SD2024
Clustering and Mining Accented Speech for Inclusive and Fair Speech Recognition
Jaeyoung Kim, Han Lu, Soheil Khorram +3
Modern automatic speech recognition (ASR) systems are typically trained on more than tens of thousands hours of speech data, which is one of the main factors for their great succes…