3 papers
eess.AS2026
Evaluating Pretrained General-Purpose Audio Representations for Music Genre Classification
Kashish Rai, Mrinmoy Bhattacharjee
This study investigates the use of self-supervised learning embeddings, particularly BYOL-A, in conjunction with a deep neural network classifier for Music Genre Classification. Ou…
eess.AS2026
Timbre-Aware LLM-based Direct Speech-to-Speech Translation Extendable to Multiple Language Pairs
Lalaram Arya, Mrinmoy Bhattacharjee, Adarsh C. R. +1
Direct Speech-to-Speech Translation (S2ST) has gained increasing attention for its ability to translate speech from one language to another, while reducing error propagation and la…
eess.AS2025
Fusion of Modulation Spectrogram and SSL with Multi-head Attention for Fake Speech Detection
Rishith Sadashiv T N, Abhishek Bedge, Saisha Suresh Bore +3
Fake speech detection systems have become a necessity to combat against speech deepfakes. Current systems exhibit poor generalizability on out-of-domain speech samples due to lack…