2 citations · 3 across the 15 of their papers we have counts for
6 papers · 1 filter
WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
Yuxue Yang, Shuyao Shang, Jiahe Wang +13
Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated…
Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding
Youze Wang, Zijun Chen, Ruoyu Chen +8
Recent advancements in multimodal large language models for video understanding (videoLLMs) have enhanced their capacity to process complex spatiotemporal data. However, challenges…
Beyond Quantity: Distribution-Aware Labeling for Visual Grounding
Yichi Zhang, Gongwei Chen, Jun Zhu +2
Visual grounding requires large and diverse region-text pairs. However, manual annotation is costly and fixed vocabularies restrict scalability and generalization. Existing pseudo-…
Effective Black-Box Multi-Faceted Attacks Breach Vision Large Language Model Guardrails
Yijun Yang, Lichao Wang, Xiao Yang +2
Vision Large Language Models (VLLMs) integrate visual data processing, expanding their real-world applications, but also increasing the risk of generating unsafe responses. In resp…
Exploring Aleatoric Uncertainty in Object Detection via Vision Foundation Models
Peng Cui, Guande He, Dan Zhang +3
Datasets collected from the open world unavoidably suffer from various forms of randomness or noiseness, leading to the ubiquity of aleatoric (data) uncertainty. Quantifying such u…
Improving transferability of 3D adversarial attacks with scale and shear transformations
Jinali Zhang, Yinpeng Dong, Jun Zhu +3
Previous work has shown that 3D point cloud classifiers can be vulnerable to adversarial examples. However, most of the existing methods are aimed at white-box attacks, where the p…