activity
20182024
most citedOn Attention Modules for Audio-Visual Synchronization

4 citations · 4 across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2023

LLM2Loss: Leveraging Language Models for Explainable Model Diagnostics

Shervin Ardeshir

Trained on a vast amount of data, Large Language models (LLMs) have achieved unprecedented success and generalization in modeling fairly complex textual inputs in the abstract spac…

cs.CV2023

Improving Identity-Robustness for Face Models

Qi Qi, Shervin Ardeshir

Despite the success of deep-learning models in many tasks, there have been concerns about such models learning shortcuts, and their lack of robustness to irrelevant confounders. Wh…

cs.CV2022

On Negative Sampling for Audio-Visual Contrastive Learning from Movies

Mahdi M. Kalayeh, Shervin Ardeshir, Lingyi Liu +2

The abundance and ease of utilizing sound, along with the fact that auditory clues reveal a plethora of information about what happens in a scene, make the audio-visual space an in…

cs.CV2022

Character-focused Video Thumbnail Retrieval

Shervin Ardeshir, Nagendra Kamath, Hossein Taghavi

We explore retrieving character-focused video frames as candidates for being video thumbnails. To evaluate each frame of the video based on the character(s) present in it, characte…

cs.CV2022

Estimating Structural Disparities for Face Models

Shervin Ardeshir, Cristina Segalin, Nathan Kallus

In machine learning, disparity metrics are often defined by measuring the difference in the performance or outcome of a model, across different sub-populations (groups) of datapoin…

cs.CV20184 cited

On Attention Modules for Audio-Visual Synchronization

Naji Khosravan, Shervin Ardeshir, Rohit Puri

With the development of media and networking technologies, multimedia applications ranging from feature presentation in a cinema setting to video on demand to interactive video con…