activity
20152024
most citedRevisiting Unreasonable Effectiveness of Data in Deep Learning Era

304 citations · 441 across the 22 of their papers we have counts for

collaborators

50 papers

cs.CV20251 cited

Mitigating Hallucinations in Diffusion Models through Adaptive Attention Modulation

Trevine Oorloff, Yaser Yacoob, Abhinav Shrivastava

Diffusion models, while increasingly adept at generating realistic images, are notably hindered by hallucinations -- unrealistic or incorrect features inconsistent with the trained…

cs.CV2024

LEIA: Latent View-invariant Embeddings for Implicit 3D Articulation

Archana Swaminathan, Anubhav Gupta, Kamal Gupta +3

Neural Radiance Fields (NeRFs) have revolutionized the reconstruction of static scenes and objects in 3D, offering unprecedented quality. However, extending NeRFs to model dynamic…

cs.CV20228 cited

CNeRV: Content-adaptive Neural Representation for Visual Data

Hao Chen, Matt Gwilliam, Bo He +2

Compression and reconstruction of visual data have been widely studied in the computer vision community, even before the popularization of deep learning. More recently, some have u…

cs.CV20228 cited

Disentangling Visual Embeddings for Attributes and Objects

Nirat Saini, Khoi Pham, Abhinav Shrivastava

We study the problem of compositional zero-shot learning for object-attribute recognition. Prior works use visual features extracted with a backbone network, pre-trained for object…

cs.CV20227 cited

LilNetX: Lightweight Networks with EXtreme Model Compression and Structured Sparsification

Sharath Girish, Kamal Gupta, Saurabh Singh +1

We introduce LilNetX, an end-to-end trainable technique for neural networks that enables learning models with specified accuracy-rate-computation trade-off. Prior works approach th…

cs.CV20225 cited

ASM-Loc: Action-aware Segment Modeling for Weakly-Supervised Temporal Action Localization

Bo He, Xitong Yang, Le Kang +3

Weakly-supervised temporal action localization aims to recognize and localize action segments in untrimmed videos given only video-level action labels for training. Without the bou…