2 papers
cs.CV2026
HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework
Jiacheng Yang, Anqi Chen, Yunkai Dang +5
Current Large Multimodal Models (LMMs) struggle with high-resolution visual inputs during the reasoning process, as the number of image tokens increases quadratically with resoluti…
cs.CR2025
FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts
Yichen Gong, Delong Ran, Jinyuan Liu +5
Large Vision-Language Models (LVLMs) signify a groundbreaking paradigm shift within the Artificial Intelligence (AI) community, extending beyond the capabilities of Large Language…