5 citations · 8 across the 22 of their papers we have counts for
17 papers · 1 filter
Do Video Generators Track the World Across Segments? A Benchmark and Method for World-State Reasoning in Video Continuation
Yingmao Miao, Pengfei Zhang, Chaoran Xu +5
Video generators build long videos by composing shorter parts, either by generating segments one after another or by autoregressively extending chunks. Each new part usually depend…
Detecting Object Hallucinations in Large Vision-Language Models via Cross-Modal Attention Drifts and Mask-Based Verification
Xuanbing Wen, Boxu Chen, Le Yang +4
Despite recent advances in large vision-language models (LVLMs), object hallucination remains a major barrier to their reliable deployment. Existing detection methods often charact…
Detecting Backdoors in Object Detection via Pre-NMS Prediction Distribution Shift
Longtian Wang, Zhengyu Zhao, Chenhao Lin +5
Object detection models deployed in safety-critical applications remain vulnerable to backdoor attacks that cause targeted misbehaviors when a hidden trigger is present. Existing d…
Evaluation-Verification Reward for Consistent Multi-Reference Image Editing
Yingmao Miao, Pengfei Zhang, Xiaochen Lv +5
While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly in maintaining visual consistency across references and ensuri…
On Success and Simplicity: A Second Look at Transferable Vision-Language Attack Pipeline
Yuchen Ren, Zhengyu Zhao, Chenhao Lin +2
Vision-Language Pre-training Models (VLPMs) are known to be vulnerable to adversarial attacks. Recent transferable attacks on VLPMs have followed a common pipeline with complicated…
Revisiting Adversarial Patch Defenses on Object Detectors: Unified Evaluation, Large-Scale Dataset, and New Insights
Junhao Zheng, Jiahao Sun, Chenhao Lin +6
Developing reliable defenses against patch attacks on object detectors has attracted increasing interest. However, we identify that existing defense evaluations lack a unified and…