5 papers
Visual Anchoring in Diffusion: Multimodal Zero-Shot Skeleton Action Recognition
Zehao Bao, Shujun Guo, Bruce X. B. Yu
Zero-shot Skeleton Action Recognition (ZSAR) remains ambiguous when unseen actions share similar skeleton joint dynamics but differ in objects or scene context. RGB provides these…
Who Decides How Knowing Becomes Doing? Redistributing Authority in Human-AI Music Co-Creation
Zhejing Hu, Yan Liu, Zhi Zhang +3
In the era of human-AI co-creation, the maxim "knowing is easy, doing is hard" is redefined. AI has the potential to ease execution, yet the essence of "hard" lies in who governs t…
MusicScaffold: Bridging Machine Efficiency and Human Growth in Adolescent Creative Education through Generative AI
Zhejing Hu, Yan Liu, Zhi Zhang +4
Adolescence is marked by strong creative impulses but limited strategies for structured expression, often leading to frustration or disengagement. While generative AI lowers techni…
CompLex: Music Theory Lexicon Constructed by Autonomous Agents for Automatic Music Generation
Zhejing Hu, Yan Liu, Gong Chen +1
Generative artificial intelligence in music has made significant strides, yet it still falls short of the substantial achievements seen in natural language processing, primarily du…
PTSM: Physiology-aware and Task-invariant Spatio-temporal Modeling for Cross-Subject EEG Decoding
Changhong Jing, Yan Liu, Shuqiang Wang +5
Cross-subject electroencephalography (EEG) decoding remains a fundamental challenge in brain-computer interface (BCI) research due to substantial inter-subject variability and the…