6 papers
When Thinking Hurts: Mitigating Visual Forgetting in Video Reasoning via Frame Repetition
Xiaokun Sun, Yubo Wang, Haoyu Cao +1
Recently, Multimodal Large Language Models (MLLMs) have demonstrated significant potential in complex visual tasks through the integration of Chain-of-Thought (CoT) reasoning. Howe…
FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing
Shida Wang, Chaohu Liu, Yubo Wang +1
Large language models represent significant investments in computation, data, and engineering expertise, making them extraordinarily valuable intellectual assets. Nevertheless, the…
CROP: Integrating Topological and Spatial Structures via Cross-View Prefixes for Molecular LLMs
Jianting Tang, Yubo Wang, Haoyu Cao +1
Recent advances in molecular science have been propelled significantly by large language models (LLMs). However, their effectiveness is limited when relying solely on molecular seq…
BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
Jianting Tang, Yubo Wang, Haoyu Cao +1
Mainstream Multimodal Large Language Models (MLLMs) achieve visual understanding by using a vision projector to bridge well-pretrained vision encoders and large language models (LL…
Tracking the Copyright of Large Vision-Language Models through Parameter Learning Adversarial Images
Yubo Wang, Jianting Tang, Chaohu Liu +1
Large vision-language models (LVLMs) have demonstrated remarkable image understanding and dialogue capabilities, allowing them to handle a variety of visual question answering task…
Break the Visual Perception: Adversarial Attacks Targeting Encoded Visual Tokens of Large Vision-Language Models
Yubo Wang, Chaohu Liu, Yanqiu Qu +3
Large vision-language models (LVLMs) integrate visual information into large language models, showcasing remarkable multi-modal conversational capabilities. However, the visual mod…