1 citations · 1 across the 6 of their papers we have counts for
5 papers · 1 filter
VisualActBench: Can VLMs See and Act like a Human?
Daoan Zhang, Pai Liu, Xiaofei Zhou +6
Vision-Language Models (VLMs) have achieved impressive progress in perceiving and describing visual environments. However, their ability to proactively reason and act based solely…
NP-LoRA: Null Space Projection for Subject-Style LoRA Fusion
Chuheng Chen, Xiaofei Zhou, Geyuan Zhang +1
Low-Rank Adaptation (LoRA) fusion enables the composition of subject and style representations for controllable generation without retraining. However, existing approaches primaril…
LaRe: Latent Refocusing for Multimodal Reasoning
Jizheng Ma, Xiaofei Zhou, Geyuan Zhang +2
Chain of Thought (CoT) reasoning enhances logical performance by decomposing complex tasks, yet its multimodal extension faces a trade-off. The prevailing Thinking with Images para…
Jailbreak Large Vision-Language Models Through Multi-Modal Linkage
Yu Wang, Xiaofei Zhou, Yichen Wang +2
With the significant advancement of Large Vision-Language Models (VLMs), concerns about their potential misuse and abuse have grown rapidly. Previous studies have highlighted VLMs'…
Exploring Explicit and Implicit Visual Relationships for Image Captioning
Zeliang Song, Xiaofei Zhou
Image captioning is one of the most challenging tasks in AI, which aims to automatically generate textual sentences for an image. Recent methods for image captioning follow encoder…