3 papers
cs.CV2025
A Benchmark for Crime Surveillance Video Analysis with Large Models
Haoran Chen, Dong Yi, Moyan Cao +3
Anomaly analysis in surveillance videos is a crucial topic in computer vision. In recent years, multimodal large language models (MLLMs) have outperformed task-specific models in v…
cs.CL2025
MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark
Dongyi Yi, Guibo Zhu, Chenglin Ding +3
With the rapid advancement of Multimodal Large Language Models (MLLMs), numerous evaluation benchmarks have emerged. However, comprehensive assessments of their performance across…
cs.CV2024
Auto DragGAN: Editing the Generative Image Manifold in an Autoregressive Manner
Pengxiang Cai, Zhiwei Liu, Guibo Zhu +2
Pixel-level fine-grained image editing remains an open challenge. Previous works fail to achieve an ideal trade-off between control granularity and inference speed. They either fai…