22 citations · 25 across the 3 of their papers we have counts for
3 papers
cs.CV2023★ 3 cited
DiffAVA: Personalized Text-to-Audio Generation with Visual Alignment
Shentong Mo, Jing Shi, Yapeng Tian
Text-to-audio (TTA) generation is a recent popular problem that aims to synthesize general audio given text descriptions. Previous methods utilized latent diffusion models to learn…
cs.CL2023★ 22 cited
X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages
Feilong Chen, Minglun Han, Haozhi Zhao +4
Large language models (LLMs) have demonstrated remarkable language abilities. GPT-4, based on advanced LLMs, exhibits extraordinary multimodal capabilities beyond previous visual l…
cs.AI2023
Mixture of personality improved Spiking actor network for efficient multi-agent cooperation
Xiyun Li, Ziyi Ni, Jingqing Ruan +4
Adaptive human-agent and agent-agent cooperation are becoming more and more critical in the research area of multi-agent reinforcement learning (MARL), where remarked progress has…