14 citations · 26 across the 3 of their papers we have counts for
5 papers
Populate-A-Scene: Affordance-Aware Human Video Generation
Mengyi Shan, Zecheng He, Haoyu Ma +4
Can a video generation model be repurposed as an interactive world simulator? We explore the affordance perception potential of text-to-video models by teaching them to predict hum…
A Cross-Verification Approach for Protecting World Leaders from Fake and Tampered Audio
Mengyi Shan, TJ Tsai
This paper tackles the problem of verifying the authenticity of speech recordings from world leaders. Whereas previous work on detecting deep fake or tampered audio focus on scruti…
Improved Handling of Repeats and Jumps in Audio-Sheet Image Synchronization
Mengyi Shan, TJ Tsai
This paper studies the problem of automatically generating piano score following videos given an audio recording and raw sheet music images. Whereas previous works focus on synthet…
Using Cell Phone Pictures of Sheet Music To Retrieve MIDI Passages
TJ Tsai, Daniel Yang, Mengyi Shan +2
This article investigates a cross-modal retrieval problem in which a user would like to retrieve a passage of music from a MIDI file by taking a cell phone picture of several lines…
MIDI Passage Retrieval Using Cell Phone Pictures of Sheet Music
Daniel Yang, Thitaree Tanprasert, Teerapat Jenrungrot +2
This paper investigates a cross-modal retrieval problem in which a user would like to retrieve a passage of music from a MIDI file by taking a cell phone picture of a physical page…