1 paper
Lin Zhang, Shentong Mo, Yijing Zhang +1
Current visual generation methods can produce high quality videos guided by texts. However, effectively controlling object dynamics remains a challenge. This work explores audio as…