activity
20162026
most citedA Lip Sync Expert Is All You Need for Speech to Lip Generation In The Wild

850 citations · 1.1k across the 56 of their papers we have counts for

collaborators
Showing 2023Show all

7 papers · 1 filter

cs.CV2023★ 1 cited

Understanding the Vulnerability of CLIP to Image Compression

Cangxiong Chen, Vinay P. Namboodiri, Julian Padget

CLIP is a widely used foundational vision-language model that is used for zero-shot image recognition and other image-text alignment tasks. We demonstrate that CLIP is vulnerable t…

cs.LG2023★ 1 cited

VERSE: Virtual-Gradient Aware Streaming Lifelong Learning with Anytime Inference

Soumya Banerjee, Vinay K. Verma, Avideep Mukherjee +3

Lifelong learning or continual learning is the problem of training an AI agent continuously while also preventing it from forgetting its previously acquired knowledge. Streaming li…

cs.LG2023

PEAR: Primitive Enabled Adaptive Relabeling for Boosting Hierarchical Reinforcement Learning

Utsav Singh, Vinay P. Namboodiri

Hierarchical reinforcement learning (HRL) has the potential to solve complex long horizon tasks using temporal abstraction and increased exploration. However, hierarchical agents a…

cs.CV2023★ 53 cited

SpectFormer: Frequency and Attention is what you need in a Vision Transformer

Badri N. Patro, Vinay P. Namboodiri, Vijay Srinivas Agneeswaran

Vision transformers have been applied successfully for image recognition tasks. There have been either multi-headed self-attention based (ViT \cite{dosovitskiy2020image}, DeIT, \ci…

cs.LG2023

CRISP: Curriculum Inducing Primitive Informed Subgoal Prediction for Hierarchical Reinforcement Learning

Utsav Singh, Vinay P. Namboodiri

Hierarchical reinforcement learning (HRL) leverages temporal abstraction to efficiently tackle complex long-horizon tasks. However, HRL often collapses because the continual update…

cs.CV2023★ 3 cited

READ Avatars: Realistic Emotion-controllable Audio Driven Avatars

Jack Saunders, Vinay Namboodiri

We present READ Avatars, a 3D-based approach for generating 2D avatars that are driven by audio input with direct and granular control over the emotion. Previous methods are unable…