Showing cs.SDShow all
2 papers · 1 filter
cs.SD2026
wav2tok 2.0: Scalable Audio Tokenization Maintaining Explicit Pairwise Token Alignment for Efficient Audio Retrieval
Adhiraj Banerjee, Vipul Arora
Learning discrete speech representations that preserve similarity across variable-length utterances is central to query-by-example spoken term detection (QbE-STD). While wav2tok in…
cs.SD2025
Continual Learning for Singing Voice Separation with Human in the Loop Adaptation
Ankur Gupta, Anshul Rai, Archit Bansal +1
Deep learning-based works for singing voice separation have performed exceptionally well in the recent past. However, most of these works do not focus on allowing users to interact…