397 citations · 991 across the 72 of their papers we have counts for
6 papers · 1 filter
Cosmos 3: Omnimodal World Models for Physical AI
NVIDIA, :, Aditi +293
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…
VILA: On Pre-training for Visual Language Models
Ji Lin, Hongxu Yin, Wei Ping +7
Visual language models (VLMs) rapidly progressed with the recent success of large language models. There have been growing efforts on visual instruction tuning to extend the LLM wi…
Re-ViLM: Retrieval-Augmented Visual Language Model for Zero and Few-Shot Image Captioning
Zhuolin Yang, Wei Ping, Zihan Liu +13
Augmenting pretrained language models (LMs) with a vision encoder (e.g., Flamingo) has obtained the state-of-the-art results in image-to-text generation. However, these models stor…
Long-Short Transformer: Efficient Transformers for Language and Vision
Chen Zhu, Wei Ping, Chaowei Xiao +4
Transformers have achieved success in both language and vision domains. However, it is prohibitively expensive to scale them to long sequences such as long documents or high-resolu…
Neural ODEs for Image Segmentation with Level Sets
Rafael Valle, Fitsum Reda, Mohammad Shoeybi +3
We propose a novel approach for image segmentation that combines Neural Ordinary Differential Equations (NODEs) and the Level Set method. Our approach parametrizes the evolution of…
Unsupervised Video Interpolation Using Cycle Consistency
Fitsum A. Reda, Deqing Sun, Aysegul Dundar +6
Learning to synthesize high frame rate videos via interpolation requires large quantities of high frame rate training videos, which, however, are scarce, especially at high resolut…