1 citations · 1 across the 4 of their papers we have counts for
4 papers · 1 filter
LaRe: Latent Refocusing for Multimodal Reasoning
Jizheng Ma, Xiaofei Zhou, Geyuan Zhang +2
Chain of Thought (CoT) reasoning enhances logical performance by decomposing complex tasks, yet its multimodal extension faces a trade-off. The prevailing Thinking with Images para…
NP-LoRA: Null Space Projection for Subject-Style LoRA Fusion
Chuheng Chen, Xiaofei Zhou, Geyuan Zhang +1
Low-Rank Adaptation (LoRA) fusion enables the composition of subject and style representations for controllable generation without retraining. However, existing approaches primaril…
VisualActBench: Can VLMs See and Act like a Human?
Daoan Zhang, Pai Liu, Xiaofei Zhou +6
Vision-Language Models (VLMs) have achieved impressive progress in perceiving and describing visual environments. However, their ability to proactively reason and act based solely…
Jailbreak Large Vision-Language Models Through Multi-Modal Linkage
Yu Wang, Xiaofei Zhou, Yichen Wang +2
With the significant advancement of Large Vision-Language Models (VLMs), concerns about their potential misuse and abuse have grown rapidly. Previous studies have highlighted VLMs'…