most citedGraph Attention for Automated Audio Captioning

14 citations · 15 across the 7 of their papers we have counts for

collaborators

7 papers

cs.SD2025

DualMark: Identifying Model and Training Data Origins in Generated Audio

Xuefeng Yang, Jian Guan, Feiyang Xiao +5

Existing watermarking methods for audio generative models only enable model-level attribution, allowing the identification of the originating generation model, but are unable to tr…

eess.IV2025

Aneumo: A Large-Scale Multimodal Aneurysm Dataset with Computational Fluid Dynamics Simulations and Deep Learning Benchmarks

Xigui Li, Yuanye Zhou, Feiyang Xiao +16

Intracranial aneurysms (IAs) are serious cerebrovascular lesions found in approximately 5\% of the general population. Their rupture may lead to high mortality. Current methods for…

cs.SD2023

Transformer-based Autoencoder with ID Constraint for Unsupervised Anomalous Sound Detection

Jian Guan, Youde Liu, Qiuqiang Kong +4

Unsupervised anomalous sound detection (ASD) aims to detect unknown anomalous sounds of devices when only normal sound data is available. The autoencoder (AE) and self-supervised l…

cs.SD20231 cited

Synth-AC: Enhancing Audio Captioning with Synthetic Supervision

Feiyang Xiao, Qiaoxi Zhu, Jian Guan +4

Data-driven approaches hold promise for audio captioning. However, the development of audio captioning methods can be biased due to the limited availability and quality of text-aud…

cs.SD2023

Anomalous Sound Detection Using Self-Attention-Based Frequency Pattern Analysis of Machine Sounds

Hejing Zhang, Jian Guan, Qiaoxi Zhu +2

Different machines can exhibit diverse frequency patterns in their emitted sound. This feature has been recently explored in anomaly sound detection and reached state-of-the-art pe…

cs.SD2023

Anomalous Sound Detection using Audio Representation with Machine ID based Contrastive Learning Pretraining

Jian Guan, Feiyang Xiao, Youde Liu +2

Existing contrastive learning methods for anomalous sound detection refine the audio representation of each audio sample by using the contrast between the samples' augmentations (e…