6 papers
KFC-KWS: Keyframe Fusion with CTC for User-Defined Keyword Spotting
Jin Li, Wenbin Jiang, Ji Hu
User-defined keyword spotting (KWS) enables personalized voice interaction by detecting user-specified keywords. A key challenge in this task is distinguishing target keywords from…
Speed by Simplicity: A Single-Stream Architecture for Fast Audio-Video Generative Foundation Model
SII-GAIR, Sand. ai, : +43
We present daVinci-MagiHuman, an open-source audio-video generative foundation model for human-centric generation. daVinci-MagiHuman jointly generates synchronized video and audio…
Adapting Speech Foundation Models for Unified Multimodal Speech Recognition with Large Language Models
Jing-Xuan Zhang, Genshun Wan, Jin Li +3
While speech foundation models (SFMs) have demonstrated remarkable performance in audio-only tasks, their adaptation to multimodal scenarios remains underexplored. This work presen…
CATP: Cross-Attention Token Pruning for Accuracy Preserved Multimodal Model Inference
Ruqi Liao, Chuqing Zhao, Jin Li +4
In response to the rising interest in large multimodal models, we introduce Cross-Attention Token Pruning (CATP), a precision-focused token pruning method. Our approach leverages c…
BUT Systems for WildSpoof Challenge: SASV in the Wild
Junyi Peng, Jin Li, Johan Rohdin +3
This paper presents the BUT submission to the WildSpoof Challenge, focusing on the Spoofing-robust Automatic Speaker Verification (SASV) track. We propose a SASV framework designed…
BUT Systems for Environmental Sound Deepfake Detection in the ESDD 2026 Challenge
Junyi Peng, Lin Zhang, Jin Li +2
This paper describes the BUT submission to the ESDD 2026 Challenge, specifically focusing on Track 1: Environmental Sound Deepfake Detection with Unseen Generators. To address the…