3 papers
cs.CV2025
UniVid: The Open-Source Unified Video Model
Jiabin Luo, Junhui Lin, Zeyu Zhang +4
Unified video modeling that combines generation and understanding capabilities is increasingly important but faces two key challenges: maintaining semantic faithfulness during flow…
cs.CV2025
VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery
Jinchao Ge, Tengfei Cheng, Biao Wu +7
Understanding cultural heritage artifacts such as ancient Greek pottery requires expert-level reasoning that remains challenging for current MLLMs due to limited domain-specific da…
cs.CV2025
PresentAgent: Multimodal Agent for Presentation Video Generation
Jingwei Shi, Zeyu Zhang, Biao Wu +4
We present PresentAgent, a multimodal agent that transforms long-form documents into narrated presentation videos. While existing approaches are limited to generating static slides…