activity
20162025
most citedFoodX-251: A Dataset for Fine-grained Food Classification

66 citations · 119 across the 11 of their papers we have counts for

collaborators
Showing cs.CVShow all

16 papers · 1 filter

cs.CV2023

A Video is Worth 10,000 Words: Training and Benchmarking with Diverse Captions for Better Long Video Retrieval

Matthew Gwilliam, Michael Cogswell, Meng Ye +3

Existing long video retrieval systems are trained and tested in the paragraph-to-video retrieval regime, where every long video is described by a single long paragraph. This neglec…

cs.CV2023

DRESS: Instructing Large Vision-Language Models to Align and Interact with Humans via Natural Language Feedback

Yangyi Chen, Karan Sikka, Michael Cogswell +2

We present DRESS, a large vision language model (LVLM) that innovatively exploits Natural Language feedback (NLF) from Large Language Models to enhance its alignment and interactio…

cs.CV2023

TIJO: Trigger Inversion with Joint Optimization for Defending Multimodal Backdoored Models

Indranil Sur, Karan Sikka, Matthew Walmer +5

We present a Multimodal Backdoor Defense technique TIJO (Trigger Inversion using Joint Optimization). Recent work arXiv:2112.07668 has demonstrated successful backdoor attacks on m…

cs.CV2021

Challenges in Procedural Multimodal Machine Comprehension:A Novel Way To Benchmark

Pritish Sahu, Karan Sikka, Ajay Divakaran

We focus on Multimodal Machine Reading Comprehension (M3C) where a model is expected to answer questions based on given passage (or context), and the context and the questions can…

cs.CV20205 cited

Zero-Shot Learning with Knowledge Enhanced Visual Semantic Embeddings

Karan Sikka, Jihua Huang, Andrew Silberfarb +6

We improve zero-shot learning (ZSL) by incorporating common-sense knowledge in DNNs. We propose Common-Sense based Neuro-Symbolic Loss (CSNL) that formulates prior knowledge as nov…

cs.CV202021 cited

RGB2LIDAR: Towards Solving Large-Scale Cross-Modal Visual Localization

Niluthpol Chowdhury Mithun, Karan Sikka, Han-Pang Chiu +2

We study an important, yet largely unexplored problem of large-scale cross-modal visual localization by matching ground RGB images to a geo-referenced aerial LIDAR 3D point cloud (…