activity
20142023
most citedDenseNet: Implementing Efficient ConvNet Descriptor Pyramids

659 citations · 792 across the 9 of their papers we have counts for

collaborators

9 papers

cs.SD20231 cited

On The Open Prompt Challenge In Conditional Audio Generation

Ernie Chang, Sidd Srinivasan, Mahi Luthra +8

Text-to-audio generation (TTA) produces audio from a text description, learning from pairs of audio samples and hand-annotated text. However, commercializing audio generation is ch…

cs.SD20231 cited

In-Context Prompt Editing For Conditional Audio Generation

Ernie Chang, Pin-Jie Lin, Yang Li +6

Distributional shift is a central challenge in the deployment of machine learning models as they can be ill-equipped for real-world data. This is particularly evident in text-to-au…

cs.CV20236 cited

CVPR 2023 Text Guided Video Editing Competition

Jay Zhangjie Wu, Xiuyu Li, Difei Gao +17

Humans watch more than a billion hours of video per day. Most of this video was edited manually, which is a tedious process. However, AI-enabled video-generation and video-editing…

eess.AS2023

Stack-and-Delay: a new codebook pattern for music generation

Gael Le Lan, Varun Nagaraja, Ernie Chang +5

In language modeling based music generation, a generated waveform is represented by a sequence of hierarchical token stacks that can be decoded either in an auto-regressive manner…

cs.SD2023

Enhance audio generation controllability through representation similarity regularization

Yangyang Shi, Gael Le Lan, Varun Nagaraja +6

This paper presents an innovative approach to enhance control over audio generation by emphasizing the alignment between audio and text representations during model training. In th…

cs.CV201614 cited

Exploring the Design Space of Deep Convolutional Neural Networks at Large Scale

Forrest Iandola

In recent years, the research community has discovered that deep neural networks (DNNs) and convolutional neural networks (CNNs) can yield higher accuracy than all previous solutio…