1 paper
Ivan Mikheev, Viacheslav Vasilev, Anna Dmitrienko +4
Text-to-audio-video (T2AV) generation models produce a video and its soundtrack from a textual description, but offer no control over whose voice speaks in the output. We show that…