collaborators

6 papers

cs.SD2026

KFC-KWS: Keyframe Fusion with CTC for User-Defined Keyword Spotting

Jin Li, Wenbin Jiang, Ji Hu

User-defined keyword spotting (KWS) enables personalized voice interaction by detecting user-specified keywords. A key challenge in this task is distinguishing target keywords from…

cs.CV2026

Speed by Simplicity: A Single-Stream Architecture for Fast Audio-Video Generative Foundation Model

SII-GAIR, Sand. ai, : +43

We present daVinci-MagiHuman, an open-source audio-video generative foundation model for human-centric generation. daVinci-MagiHuman jointly generates synchronized video and audio…

eess.AS2026

Adapting Speech Foundation Models for Unified Multimodal Speech Recognition with Large Language Models

Jing-Xuan Zhang, Genshun Wan, Jin Li +3

While speech foundation models (SFMs) have demonstrated remarkable performance in audio-only tasks, their adaptation to multimodal scenarios remains underexplored. This work presen…

cs.CL2026

CATP: Cross-Attention Token Pruning for Accuracy Preserved Multimodal Model Inference

Ruqi Liao, Chuqing Zhao, Jin Li +4

In response to the rising interest in large multimodal models, we introduce Cross-Attention Token Pruning (CATP), a precision-focused token pruning method. Our approach leverages c…

eess.AS2025

BUT Systems for WildSpoof Challenge: SASV in the Wild

Junyi Peng, Jin Li, Johan Rohdin +3

This paper presents the BUT submission to the WildSpoof Challenge, focusing on the Spoofing-robust Automatic Speaker Verification (SASV) track. We propose a SASV framework designed…

eess.AS2025

BUT Systems for Environmental Sound Deepfake Detection in the ESDD 2026 Challenge

Junyi Peng, Lin Zhang, Jin Li +2

This paper describes the BUT submission to the ESDD 2026 Challenge, specifically focusing on Track 1: Environmental Sound Deepfake Detection with Unseen Generators. To address the…