21 citations · 21 across the 3 of their papers we have counts for
3 papers · 1 filter
CRoPS: A Training-Free Hallucination Mitigation Framework for Vision-Language Models
Neeraj Anand, Samyak Jha, Udbhav Bamba +1
Despite the rapid success of Large Vision-Language Models (LVLMs), a persistent challenge is their tendency to generate hallucinated content, undermining reliability in real-world…
Coarse to Fine Multi-Resolution Temporal Convolutional Network
Dipika Singhania, Rahul Rahaman, Angela Yao
Temporal convolutional networks (TCNs) are a commonly used architecture for temporal video segmentation. TCNs however, tend to suffer from over-segmentation errors and require addi…
Pretrained equivariant features improve unsupervised landmark discovery
Rahul Rahaman, Atin Ghosh, Alexandre H. Thiery
Locating semantically meaningful landmark points is a crucial component of a large number of computer vision pipelines. Because of the small number of available datasets with groun…