5 citations · 9 across the 5 of their papers we have counts for
5 papers
GPT-4V(ision) as A Social Media Analysis Engine
Hanjia Lyu, Jinfa Huang, Daoan Zhang +6
Recent research has offered insights into the extraordinary capabilities of Large Multimodal Models (LMMs) in various general vision and language tasks. There is growing interest i…
MobileVidFactory: Automatic Diffusion-Based Social Media Video Generation for Mobile Devices from Text
Junchen Zhu, Huan Yang, Wenjing Wang +8
Videos for mobile devices become the most popular access to share and acquire information recently. For the convenience of users' creation, in this paper, we present a system, name…
Deficiency-Aware Masked Transformer for Video Inpainting
Yongsheng Yu, Heng Fan, Libo Zhang
Recent video inpainting methods have made remarkable progress by utilizing explicit guidance, such as optical flow, to propagate cross-frame pixels. However, there are cases where…
High-Fidelity Image Inpainting with GAN Inversion
Yongsheng Yu, Libo Zhang, Heng Fan +1
Image inpainting seeks a semantically consistent way to recover the corrupted image in the light of its unmasked content. Previous approaches usually reuse the well-trained GAN as…
Unbiased Multi-Modality Guidance for Image Inpainting
Yongsheng Yu, Dawei Du, Libo Zhang +1
Image inpainting is an ill-posed problem to recover missing or damaged image content based on incomplete images with masks. Previous works usually predict the auxiliary structures…