2 papers
cs.GR2026
DrawVideo: Generating Long Video from Storyboard Keyframe Sketches
Chuanzhi Xu, Huiqi Liang, Bang Shi +7
Long video generation requires high-fidelity synthesis, coherent narrative structure, and user control over extended time spans. Existing text-to-video methods often rely on a sing…
eess.AS2026
A Survey of Audio Reasoning in Multimodal Foundation Models
Zhihan Guo, Wenqian Cui, Guan-Ting Lin +8
Reasoning has become a defining capability of modern foundation models, yet its development in the audio modality remains limited. Audio poses challenges that are distinct from tho…