2 papers
cs.SD2026
Multi-Task Multi-Frame Visual Piano Transcription
Yonghyun Kim, Hoyeol Sohn, Juhan Nam +1
Audio-based piano transcription performs well on onset, pitch, and velocity, but the sustain pedal lets sound persist long after key release, so audio systems predict pedal-extende…
eess.AS2026
DTM-Codec: Dynamic Token Masking for VFR Speech Coding with Efficient Boundary Selection
Hoyeol Sohn, Juhan Nam
Variable frame rate (VFR) coding has recently emerged in neural speech codecs, allocating fewer frames to redundant regions and more frames to rapidly changing speech. VFR must tra…