290 citations · 448 across the 29 of their papers we have counts for
23 papers · 1 filter
X-Fake: Juggling Utility Evaluation and Explanation of Simulated SAR Images
Zhongling Huang, Yihan Zhuang, Zipei Zhong +3
SAR image simulation has attracted much attention due to its great potential to supplement the scarce training data for deep learning algorithms. Consequently, evaluating the quali…
VDG: Vision-Only Dynamic Gaussian for Driving Simulation
Hao Li, Jingfeng Li, Dingwen Zhang +7
Dynamic Gaussian splatting has led to impressive scene reconstruction and image synthesis advances in novel views. Existing methods, however, heavily rely on pre-computed poses and…
PVUW 2024 Challenge on Complex Video Understanding: Methods and Results
Henghui Ding, Chang Liu, Yunchao Wei +34
Pixel-level Video Understanding in the Wild Challenge (PVUW) focus on complex video understanding. In this CVPR 2024 workshop, we add two new tracks, Complex Video Object Segmentat…
Energy-Latency Manipulation of Multi-modal Large Language Models via Verbose Samples
Kuofeng Gao, Jindong Gu, Yang Bai +4
Despite the exceptional performance of multi-modal large language models (MLLMs), their deployment requires substantial computational resources. Once malicious users induce high en…
AnimateZoo: Zero-shot Video Generation of Cross-Species Animation via Subject Alignment
Yuanfeng Xu, Yuhao Chen, Zhongzhan Huang +4
Recent video editing advancements rely on accurate pose sequences to animate subjects. However, these efforts are not suitable for cross-species animation due to pose misalignment…
RoDLA: Benchmarking the Robustness of Document Layout Analysis Models
Yufan Chen, Jiaming Zhang, Kunyu Peng +4
Before developing a Document Layout Analysis (DLA) model in real-world applications, conducting comprehensive robustness testing is essential. However, the robustness of DLA models…