16 citations · 27 across the 17 of their papers we have counts for
17 papers
MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding
Bo He, Hengduo Li, Young Kyun Jang +5
With the success of large language models (LLMs), integrating the vision model into LLMs to build vision-language foundation models has gained much more interest recently. However,…
What is Point Supervision Worth in Video Instance Segmentation?
Shuaiyi Huang, De-An Huang, Zhiding Yu +5
Video instance segmentation (VIS) is a challenging vision task that aims to detect, segment, and track objects in videos. Conventional VIS methods rely on densely-annotated object…
Measuring Style Similarity in Diffusion Models
Gowthami Somepalli, Anubhav Gupta, Kamal Gupta +5
Generative models are now widely used by graphic designers and artists. Prior works have shown that these models remember and often replicate content from their training data durin…
SHACIRA: Scalable HAsh-grid Compression for Implicit Neural Representations
Sharath Girish, Abhinav Shrivastava, Kamal Gupta
Implicit Neural Representations (INR) or neural fields have emerged as a popular framework to encode multimedia signals such as images and radiance fields while retaining high-qual…
Chop & Learn: Recognizing and Generating Object-State Compositions
Nirat Saini, Hanyu Wang, Archana Swaminathan +4
Recognizing and generating object-state compositions has been a challenging task, especially when generalizing to unseen compositions. In this paper, we study the task of cutting o…
Diff2Lip: Audio Conditioned Diffusion Models for Lip-Synchronization
Soumik Mukhopadhyay, Saksham Suri, Ravi Teja Gadde +1
The task of lip synchronization (lip-sync) seeks to match the lips of human faces with different audio. It has various applications in the film industry as well as for creating vir…