5 papers
VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery
Jinchao Ge, Tengfei Cheng, Biao Wu +7
Understanding cultural heritage artifacts such as ancient Greek pottery requires expert-level reasoning that remains challenging for current MLLMs due to limited domain-specific da…
UniVid: The Open-Source Unified Video Model
Jiabin Luo, Junhui Lin, Zeyu Zhang +4
Unified video modeling that combines generation and understanding capabilities is increasingly important but faces two key challenges: maintaining semantic faithfulness during flow…
ESA: Annotation-Efficient Active Learning for Semantic Segmentation
Jinchao Ge, Zeyu Zhang, Minh Hieu Phan +4
Active learning enhances annotation efficiency by selecting the most revealing samples for labeling, thereby reducing reliance on extensive human input. Previous methods in semanti…
PresentAgent: Multimodal Agent for Presentation Video Generation
Jingwei Shi, Zeyu Zhang, Biao Wu +4
We present PresentAgent, a multimodal agent that transforms long-form documents into narrated presentation videos. While existing approaches are limited to generating static slides…
MMCLIP: Cross-modal Attention Masked Modelling for Medical Language-Image Pre-Training
Biao Wu, Yutong Xie, Zeyu Zhang +4
Vision-and-language pretraining (VLP) in the medical field utilizes contrastive learning on image-text pairs to achieve effective transfer across tasks. Yet, current VLP approaches…