activity
20162026
most citedA Lip Sync Expert Is All You Need for Speech to Lip Generation In The Wild

850 citations · 1.1k across the 39 of their papers we have counts for

collaborators

74 papers

cs.CV2026

A Plug-in Interpretation of Conditioning in Score-Based Diffusion Models

Libo Chen, Souvik Ghosh, Teo Deveney +2

We propose a conditioning mechanism for diffusion models based on multi-speed joint diffusion of the target and the condition. The mechanism learns an unconditional joint score net…

cs.CV2026

Can Unsupervised Segmentation Reduce Annotation Costs for Video Semantic Segmentation?

Samik Some, Vinay P. Namboodiri

Present-day deep neural networks for video semantic segmentation require a large number of fine-grained pixel-level annotations to achieve the best possible results. Obtaining such…

cs.CV2025

EIDT-V: Exploiting Intersections in Diffusion Trajectories for Model-Agnostic, Zero-Shot, Training-Free Text-to-Video Generation

Diljeet Jagpal, Xi Chen, Vinay P. Namboodiri

Zero-shot, training-free, image-based text-to-video generation is an emerging area that aims to generate videos using existing image-based diffusion models. Current methods in this…

cs.CV2024

Self-supervised Representation Learning for Cell Event Recognition through Time Arrow Prediction

Cangxiong Chen, Vinay P. Namboodiri, Julia E. Sero

The spatio-temporal nature of live-cell microscopy data poses challenges in the analysis of cell states which is fundamental in bioimaging. Deep-learning based segmentation or trac…

cs.CV2024

RISSOLE: Parameter-efficient Diffusion Models via Block-wise Generation and Retrieval-Guidance

Avideep Mukherjee, Soumya Banerjee, Piyush Rai +1

Diffusion-based models demonstrate impressive generation capabilities. However, they also have a massive number of parameters, resulting in enormous model sizes, thus making them u…

cs.CV20231 cited

Understanding the Vulnerability of CLIP to Image Compression

Cangxiong Chen, Vinay P. Namboodiri, Julian Padget

CLIP is a widely used foundational vision-language model that is used for zero-shot image recognition and other image-text alignment tasks. We demonstrate that CLIP is vulnerable t…