8 papers
Patch as Node: Human-Centric Graph Representation Learning for Multimodal Action Recognition
Zeyu Liang, Hailun Xia, Naichuan Zheng
While human action recognition has witnessed notable achievements, multimodal methods fusing RGB and skeleton modalities still suffer from their inherent heterogeneity and fail to…
SNN-Driven Multimodal Human Action Recognition via Sparse Spatial-Temporal Data Fusion
Naichuan Zheng, Hailun Xia, Zeyu Liang +1
Multimodal human action recognition based on RGB and skeleton data fusion, while effective, is constrained by significant limitations such as high computational complexity, excessi…
LLM Collaboration With Multi-Agent Reinforcement Learning
Shuo Liu, Tianle Chen, Zeyu Liang +2
A large amount of work has been done in Multi-Agent Systems (MAS) for modeling and solving problems with multiple interacting agents. However, most LLMs are pretrained independentl…
MK-SGN: A Spiking Graph Convolutional Network with Multimodal Fusion and Knowledge Distillation for Skeleton-based Action Recognition
Naichuan Zheng, Hailun Xia, Zeyu Liang +1
In recent years, multimodal Graph Convolutional Networks (GCNs) have achieved remarkable performance in skeleton-based action recognition. The reliance on high-energy-consuming con…
Signal-SGN: A Spiking Graph Convolutional Network for Skeletal Action Recognition via Learning Temporal-Frequency Dynamics
Naichuan Zheng, Yuchen Du, Hailun Xia +1
For multimodal skeleton-based action recognition, Graph Convolutional Networks (GCNs) are effective models. Still, their reliance on floating-point computations leads to high energ…
Empowering Global Voices: A Data-Efficient, Phoneme-Tone Adaptive Approach to High-Fidelity Speech Synthesis
Yizhong Geng, Jizhuo Xu, Zeyu Liang +3
Text-to-speech (TTS) technology has achieved impressive results for widely spoken languages, yet many under-resourced languages remain challenged by limited data and linguistic com…