3 papers
eess.AS2026
Time-Layer Adaptive Alignment for Speaker Similarity in Flow-Matching Based Zero-Shot TTS
Haoyu Li, Mingyang Han, Yu Xi +8
Flow-Matching (FM)-based zero-shot text-to-speech (TTS) systems exhibit high-quality speech synthesis and robust generalization capabilities. However, the speaker representation ab…
cs.CV2026
StreamSense: Streaming Social Task Detection with Selective Vision-Language Model Routing
Han Wang, Deyi Ji, Lanyun Zhu +2
Live streaming platforms require real-time monitoring and reaction to social signals, utilizing partial and asynchronous evidence from video, text, and audio. We propose StreamSens…
cs.LG2025
Multi-Agent VLMs Guided Self-Training with PNU Loss for Low-Resource Offensive Content Detection
Han Wang, Deyi Ji, Junyu Lu +6
Accurate detection of offensive content on social media demands high-quality labeled data; however, such data is often scarce due to the low prevalence of offensive instances and t…