2 papers
cs.CV2024
Efficient Diffusion Models: A Comprehensive Survey from Principles to Practices
Zhiyuan Ma, Yuzhu Zhang, Guoli Jia +7
As one of the most popular and sought-after generative models in the recent years, diffusion models have sparked the interests of many researchers and steadily shown excellent adva…
cs.CV2024
EventLens: Leveraging Event-Aware Pretraining and Cross-modal Linking Enhances Visual Commonsense Reasoning
Mingjie Ma, Zhihuan Yu, Yichao Ma +1
Visual Commonsense Reasoning (VCR) is a cognitive task, challenging models to answer visual questions requiring human commonsense, and to provide rationales explaining why the answ…