most citedConchShell: A Generative Adversarial Networks that Turns Pictures into Piano Music

1 citations · 1 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL2025

EulerESG: Automating ESG Disclosure Analysis with LLMs

Yi Ding, Xushuo Tang, Zhengyi Yang +12

Environmental, Social, and Governance (ESG) reports have become central to how companies communicate climate risk, social impact, and governance practices, yet they are still publi…

cs.RO2025

ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver

Wenxuan Song, Ziyang Zhou, Han Zhao +7

Recent advances in Vision-Language-Action (VLA) models have enabled robotic agents to integrate multimodal understanding with action execution. However, our empirical analysis reve…

cs.RO2025

CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding

Wenxuan Song, Jiayi Chen, Pengxiang Ding +4

In recent years, Vision-Language-Action (VLA) models have become a vital research direction in robotics due to their impressive multimodal understanding and generalization capabili…

cs.SD20221 cited

ConchShell: A Generative Adversarial Networks that Turns Pictures into Piano Music

Wanpeng Fan, Yuanzhi Su, Yuxin Huang

We present ConchShell, a multi-modal generative adversarial framework that takes pictures as input to the network and generates piano music samples that match the picture context.…

eess.AS2022

PaddleSpeech: An Easy-to-Use All-in-One Speech Toolkit

Hui Zhang, Tian Yuan, Junkun Chen +10

PaddleSpeech is an open-source all-in-one speech toolkit. It aims at facilitating the development and research of speech processing technologies by providing an easy-to-use command…