3 citations · 5 across the 2 of their papers we have counts for
2 papers
cs.CV2024★ 2 cited
EventLens: Leveraging Event-Aware Pretraining and Cross-modal Linking Enhances Visual Commonsense Reasoning
Mingjie Ma, Zhihuan Yu, Yichao Ma +1
Visual Commonsense Reasoning (VCR) is a cognitive task, challenging models to answer visual questions requiring human commonsense, and to provide rationales explaining why the answ…
cs.CV2023★ 3 cited
LMD: Faster Image Reconstruction with Latent Masking Diffusion
Zhiyuan Ma, zhihuan yu, Jianjun Li +1
As a class of fruitful approaches, diffusion probabilistic models (DPMs) have shown excellent advantages in high-resolution image reconstruction. On the other hand, masked autoenco…