activity
20172026
most citedDeep Voice: Real-time Neural Text-to-Speech

397 citations · 991 across the 72 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

Cosmos 3: Omnimodal World Models for Physical AI

NVIDIA, :, Aditi +293

We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…

cs.CV2023★ 6 cited

VILA: On Pre-training for Visual Language Models

Ji Lin, Hongxu Yin, Wei Ping +7

Visual language models (VLMs) rapidly progressed with the recent success of large language models. There have been growing efforts on visual instruction tuning to extend the LLM wi…

cs.CV2023★ 1 cited

Re-ViLM: Retrieval-Augmented Visual Language Model for Zero and Few-Shot Image Captioning

Zhuolin Yang, Wei Ping, Zihan Liu +13

Augmenting pretrained language models (LMs) with a vision encoder (e.g., Flamingo) has obtained the state-of-the-art results in image-to-text generation. However, these models stor…

cs.CV2021

Long-Short Transformer: Efficient Transformers for Language and Vision

Chen Zhu, Wei Ping, Chaowei Xiao +4

Transformers have achieved success in both language and vision domains. However, it is prohibitively expensive to scale them to long sequences such as long documents or high-resolu…

cs.CV2019★ 4 cited

Neural ODEs for Image Segmentation with Level Sets

Rafael Valle, Fitsum Reda, Mohammad Shoeybi +3

We propose a novel approach for image segmentation that combines Neural Ordinary Differential Equations (NODEs) and the Level Set method. Our approach parametrizes the evolution of…

cs.CV2019

Unsupervised Video Interpolation Using Cycle Consistency

Fitsum A. Reda, Deqing Sun, Aysegul Dundar +6

Learning to synthesize high frame rate videos via interpolation requires large quantities of high frame rate training videos, which, however, are scarce, especially at high resolut…