4 citations · 4 across the 3 of their papers we have counts for
3 papers
cs.CV2026
Beyond Shortcuts: Mitigating Visual Illusions in Frozen VLMs via Qualitative Reasoning
Hao Guo, Fei Wang, Junjie Chen +4
While Vision-Language Models (VLMs) have achieved state-of-the-art performance in general visual tasks, their perceptual robustness remains remarkably brittle when confronted with…
cs.CV2026
3D Smoke Scene Reconstruction Guided by Vision Priors from Multimodal Large Language Models
Xinye Zheng, Fei Wang, Yiqi Nie +5
Reconstructing 3D scenes from smoke-degraded multi-view images is particularly difficult because smoke introduces strong scattering effects, view-dependent appearance changes, and…
cs.CV2026★ 4 cited
Real-Time Oriented Object Detection Transformer in Remote Sensing Images
Zeyu Ding, Yong Zhou, Jiaqi Zhao +4
Recent real-time detection transformers have gained popularity due to their simplicity and efficiency. However, these detectors do not explicitly model object rotation, especially…