102 citations · 233 across the 45 of their papers we have counts for
5 papers · 1 filter
TI2V-Zero: Zero-Shot Image Conditioning for Text-to-Video Diffusion Models
Haomiao Ni, Bernhard Egger, Suhas Lohit +5
Text-conditioned image-to-video generation (TI2V) aims to synthesize a realistic video starting from a given image (e.g., a woman's photo) and a text description (e.g., "a woman is…
SuperLoRA: Parameter-Efficient Unified Adaptation of Multi-Layer Attention Modules
Xiangyu Chen, Jing Liu, Ye Wang +4
Low-rank adaptation (LoRA) and its variants are widely employed in fine-tuning large models, including large language models for natural language processing and diffusion models fo…
Steered Diffusion: A Generalized Framework for Plug-and-Play Conditional Image Synthesis
Nithin Gopalakrishnan Nair, Anoop Cherian, Suhas Lohit +4
Conditional generative models typically demand large annotated training sets to achieve high-quality synthesis. As a result, there has been significant interest in designing models…
Multi-Image Visual Question Answering
Harsh Raj, Janhavi Dadhania, Akhilesh Bhardwaj +1
While a lot of work has been done on developing models to tackle the problem of Visual Question Answering, the ability of these models to relate the question to the image features…
LUVLi Face Alignment: Estimating Landmarks' Location, Uncertainty, and Visibility Likelihood
Abhinav Kumar, Tim K. Marks, Wenxuan Mou +6
Modern face alignment methods have become quite accurate at predicting the locations of facial landmarks, but they do not typically estimate the uncertainty of their predicted loca…