2 papers
cs.CV2026
TD3Net: A temporal densely connected multi-dilated convolutional network for lipreading
Byung Hoon Lee, Wooseok Shin, Sung Won Han
The word-level lipreading approach typically employs a two-stage framework with separate frontend and backend architectures to model dynamic lip movements. Each component has been…
eess.AS2024
Patch-Mix Contrastive Learning with Audio Spectrogram Transformer on Respiratory Sound Classification
Sangmin Bae, June-Woo Kim, Won-Yang Cho +7
Respiratory sound contains crucial information for the early diagnosis of fatal lung diseases. Since the COVID-19 pandemic, there has been a growing interest in contact-free medica…