Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations
Jianing Guo, Zhenhong Wu, Chang Tu +13
In Vision-Language-Actionf(VLA) models, robustness to real-world perturbations is critical for deployment. Existing methods target simple visual disturbances, overlooking the broad…
cs.CV2026
What, Whether and How? Unveiling Process Reward Models for Thinking with Images Reasoning
Yujin Zhou, Pengcheng Wen, Jiale Chen +6
The rapid advancement of Large Vision Language Models (LVLMs) has demonstrated excellent abilities in various visual tasks. Building upon these developments, the thinking with imag…
cs.CV2024
SafeSora: Towards Safety Alignment of Text2Video Generation via a Human Preference Dataset
Josef Dai, Tianle Chen, Xuyao Wang +4
To mitigate the risk of harmful outputs from large vision models (LVMs), we introduce the SafeSora dataset to promote research on aligning text-to-video generation with human value…