1 citations · 1 across the 3 of their papers we have counts for
3 papers
eess.AS2025
Music Tempo Estimation on Solo Instrumental Performance
Zhanhong He, Roberto Togneri, Xiangyu Zhang
Recently, automatic music transcription has made it possible to convert musical audio into accurate MIDI. However, the resulting MIDI lacks music notations such as tempo, which hin…
cs.CV2025
M-DocSum: Do LVLMs Genuinely Comprehend Interleaved Image-Text in Document Summarization?
Haolong Yan, Kaijun Tan, Yeqing Shen +5
We investigate a critical yet under-explored question in Large Vision-Language Models (LVLMs): Do LVLMs genuinely comprehend interleaved image-text in the document? Existing docume…
cs.CV2025★ 1 cited
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Guoqing Ma, Haoyang Huang, Kun Yan +112
We present Step-Video-T2V, a state-of-the-art text-to-video pre-trained model with 30B parameters and the ability to generate videos up to 204 frames in length. A deep compression…