5 citations · 7 across the 6 of their papers we have counts for
5 papers · 1 filter
ForenX: Towards Explainable AI-Generated Image Detection with Multimodal Large Language Models
Chuangchuang Tan, Jinglu Wang, Xiang Ming +4
Advances in generative models have led to AI-generated images visually indistinguishable from authentic ones. Despite numerous studies on detecting AI-generated images with classif…
SVLTA: Benchmarking Vision-Language Temporal Alignment via Synthetic Video Situation
Hao Du, Bo Wu, Yan Lu +1
Vision-language temporal alignment is a crucial capability for human dynamic recognition and cognition in real-world scenarios. While existing research focuses on capturing vision-…
Uncertainty-Aware Deep Video Compression with Ensembles
Wufei Ma, Jiahao Li, Bin Li +1
Deep learning-based video compression is a challenging task, and many previous state-of-the-art learning-based video codecs use optical flows to exploit the temporal correlation be…
Efficient View Synthesis with Neural Radiance Distribution Field
Yushuang Wu, Xiao Li, Jinglu Wang +3
Recent work on Neural Radiance Fields (NeRF) has demonstrated significant advances in high-quality view synthesis. A major limitation of NeRF is its low rendering efficiency due to…
StableVideo: Text-driven Consistency-aware Diffusion Video Editing
Wenhao Chai, Xun Guo, Gaoang Wang +1
Diffusion-based methods can generate realistic images and videos, but they struggle to edit existing objects in a video while preserving their appearance over time. This prevents d…