3 papers
cs.CV2025
OIG-Bench: A Multi-Agent Annotated Benchmark for Multimodal One-Image Guides Understanding
Jiancong Xie, Wenjin Wang, Zhuomeng Zhang +5
Recent advances in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities. However, evaluating their capacity for human-like understanding in One-Image…
cs.CV2025
Evading Data Provenance in Deep Neural Networks
Hongyu Zhu, Sichu Liang, Wenwen Wang +3
Modern over-parameterized deep models are highly data-dependent, with large scale general-purpose and domain-specific datasets serving as the bedrock for rapid advancements. Howeve…
cs.CV2025
PromptLA: Towards Integrity Verification of Black-box Text-to-Image Diffusion Models
Zhuomeng Zhang, Fangqi Li, Chong Di +3
Despite the impressive synthesis quality of text-to-image (T2I) diffusion models, their black-box deployment poses significant regulatory challenges: Malicious actors can fine-tune…