Showing cs.SDShow all
2 papers · 1 filter
cs.SD2026
Delayed Commitment for Representation Readiness in Stage-wise Audio-Visual Learning
Xinmeng Xu, Haoran Xie, S. Joe Qin +3
Stage-wise audio-visual encoders propagate fused intermediate states across layers, making the formation of later representations depend on the readiness of earlier fusion states.…
cs.SD2026
ATRIE: Adaptive Tuning for Robust Inference and Emotion in Persona-Driven Speech Synthesis
Aoduo Li, Haoran Lv, Hongjian Xu +5
High-fidelity character voice synthesis is a cornerstone of immersive multimedia applications, particularly for interacting with anime avatars and digital humans. However, existing…