activity
20162024
most citedA Recurrent Encoder-Decoder Network for Sequential Face Alignment

33 citations · 129 across the 17 of their papers we have counts for

collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV2024

Comparison Visual Instruction Tuning

Wei Lin, Muhammad Jehanzeb Mirza, Sivan Doveh +4

Comparing two images in terms of Commonalities and Differences (CaD) is a fundamental human capability that forms the basis of advanced visual reasoning and interpretation. It is e…

cs.CV20231 cited

Learning Human Action Recognition Representations Without Real Humans

Howard Zhong, Samarth Mishra, Donghyun Kim +7

Pre-training on massive video datasets has become essential to achieve high action recognition performance on smaller downstream datasets. However, most large-scale video datasets…

cs.CV2023

TAP: Targeted Prompting for Task Adaptive Generation of Textual Training Instances for Visual Classification

M. Jehanzeb Mirza, Leonid Karlinsky, Wei Lin +3

Vision and Language Models (VLMs), such as CLIP, have enabled visual recognition of a potentially unlimited set of categories described by text prompts. However, for the best visua…

cs.CV202312 cited

Dense and Aligned Captions (DAC) Promote Compositional Reasoning in VL Models

Sivan Doveh, Assaf Arbelle, Sivan Harary +9

Vision and Language (VL) models offer an effective method for aligning representation spaces of images and text, leading to numerous applications such as cross-modal retrieval, vis…

cs.CV20231 cited

Mind the Backbone: Minimizing Backbone Distortion for Robust Object Detection

Kuniaki Saito, Donghyun Kim, Piotr Teterwak +2

Building object detectors that are robust to domain shifts is critical for real-world applications. Prior approaches fine-tune a pre-trained backbone and risk overfitting it to in-…

cs.CV20214 cited

Unsupervised Domain Generalization by Learning a Bridge Across Domains

Sivan Harary, Eli Schwartz, Assaf Arbelle +11

The ability to generalize learned representations across significantly different visual domains, such as between real photos, clipart, paintings, and sketches, is a fundamental cap…