1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.CV2023★ 1 cited
ANetQA: A Large-scale Benchmark for Fine-grained Compositional Reasoning over Untrimmed Videos
Zhou Yu, Lixiang Zheng, Zhou Zhao +4
Building benchmarks to systemically analyze different capabilities of video question answering (VideoQA) models is challenging yet crucial. Existing benchmarks often use non-compos…
cs.CV2023
Multi-task Paired Masking with Alignment Modeling for Medical Vision-Language Pre-training
Ke Zhang, Yan Yang, Jun Yu +4
In recent years, the growing demand for medical imaging diagnosis has placed a significant burden on radiologists. As a solution, Medical Vision-Language Pre-training (Med-VLP) met…
cs.CV2023
GLOW: Global Layout Aware Attacks on Object Detection
Buyu Liu, BaoJun, Jianping Fan +3
Adversarial attacks aim to perturb images such that a predictor outputs incorrect results. Due to the limited research in structured attacks, imposing consistency checks on natural…