2 papers
cs.SD2026
AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing
Ziyang Ma, Zhikang Niu, Wenming Tu +30
We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context. To sup…
cs.MM2026
Vision-Guided Text Prompt Tuning for Multimodal Sentiment Analysis
Xiaoran Kou, Jingyi Wu, Peng Sun +2
Multimodal sentiment analysis requires effective modeling of both verbal semantics and non-verbal affective cues. A central challenge is to calibrate text-centered sentiment unders…