8 papers
Cosmos 3: Omnimodal World Models for Physical AI
NVIDIA, :, Aditi +293
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…
Beyond Encoder Accumulation: Measuring Encoder Roles in Multi-Encoder VLMs
Wei Ding, Yudong Zhang, Ruobing Xie +3
As foundation models scale toward fusing more heterogeneous visual streams, understanding how diverse encoders interact under joint training becomes a prerequisite for principled d…
MHSA: A Lightweight Framework for Mitigating Hallucinations via Steered Attention in LVLMs
Wei Ding, Yilin Li, Yudong Zhang +4
Large vision-language models (LVLMs) have achieved remarkable performance across diverse multimodal tasks, yet they continue to suffer from hallucinations, generating content that…
The Security Threat of Compressed Projectors in Large Vision-Language Models
Yudong Zhang, Ruobing Xie, Xingwu Sun +4
The choice of a suitable visual language projector (VLP) is critical to the successful training of large visual language models (LVLMs). Mainstream VLPs can be broadly categorized…
Fighting Fire with Fire (F3): A Training-free and Efficient Visual Adversarial Example Purification Method in LVLMs
Yudong Zhang, Ruobing Xie, Yiqing Huang +5
Recent advances in large vision-language models (LVLMs) have showcased their remarkable capabilities across a wide range of multimodal vision-language tasks. However, these models…
DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models
Yudong Zhang, Ruobing Xie, Xingwu Sun +5
Large vision-language models (LVLMs) have demonstrated exceptional performance on complex multimodal tasks. However, they continue to suffer from significant hallucination issues,…