11 citations · 13 across the 5 of their papers we have counts for
4 papers · 1 filter
VideoPoet: A Large Language Model for Zero-Shot Video Generation
Dan Kondratyuk, Lijun Yu, Xiuye Gu +28
We present VideoPoet, a language model capable of synthesizing high-quality video, with matching audio, from a large variety of conditioning signals. VideoPoet employs a decoder-on…
MM-AU:Towards Multimodal Understanding of Advertisement Videos
Digbalay Bose, Rajat Hebbar, Tiantian Feng +3
Advertisement videos (ads) play an integral part in the domain of Internet e-commerce as they amplify the reach of particular products to a broad audience or can serve as a medium…
MovieCLIP: Visual Scene Recognition in Movies
Digbalay Bose, Rajat Hebbar, Krishna Somandepalli +5
Longform media such as movies have complex narrative structures, with events spanning a rich variety of ambient visual scenes. Domain specific challenges associated with visual sce…
Understanding of Emotion Perception from Art
Digbalay Bose, Krishna Somandepalli, Souvik Kundu +3
Computational modeling of the emotions evoked by art in humans is a challenging problem because of the subjective and nuanced nature of art and affective signals. In this paper, we…