collaborators

6 papers

cs.CV2026

Cosmos 3: Omnimodal World Models for Physical AI

NVIDIA, :, Aditi +293

We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…

cs.CV2026

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing

Andong Deng, Dawei Du, Zhenfang Chen +7

Real-world video editing demands not only expert knowledge of cinematic techniques but also multimodal reasoning to select, align, and combine footage into coherent narratives. Whi…

cs.CV2026

Vidi2.5: Large Multimodal Models for Video Understanding and Creation

Vidi Team, Chia-Wen Kuo, Chuang Huang +31

Video has emerged as the primary medium for communication and creativity on the Internet, driving strong demand for scalable, high-quality video production. Vidi models continue to…

cs.CV2025

D-Attn: Decomposed Attention for Large Vision-and-Language Models

Chia-Wen Kuo, Sijie Zhu, Fan Chen +2

Large vision-and-language models (LVLMs) have traditionally integrated visual and textual tokens by concatenating them into a single homogeneous input for large language models (LL…

cs.CV2025

Vidi: Large Multimodal Models for Video Understanding and Editing

Vidi Team, Celong Liu, Chia-Wen Kuo +20

Humans naturally share information with those they are connected to, and video has become one of the dominant mediums for communication and expression on the Internet. To support t…

cs.CV2025

Where do Large Vision-Language Models Look at when Answering Questions?

Xiaoying Xing, Chia-Wen Kuo, Li Fuxin +6

Large Vision-Language Models (LVLMs) have shown promising performance in vision-language understanding and reasoning tasks. However, their visual understanding behaviors remain und…