30 citations · 74 across the 14 of their papers we have counts for
14 papers
ChartLlama: A Multimodal LLM for Chart Understanding and Generation
Yucheng Han, Chi Zhang, Xin Chen +5
Multi-modal large language models have demonstrated impressive performances on most vision-language tasks. However, the model generally lacks the understanding capabilities for spe…
Tuning Multi-mode Token-level Prompt Alignment across Modalities
Dongsheng Wang, Miaoge Li, Xinyang Liu +3
Advancements in prompt tuning of vision-language models have underscored their potential in enhancing open-world visual concept comprehension. However, prior works only primarily f…
Invariant Feature Regularization for Fair Face Recognition
Jiali Ma, Zhongqi Yue, Kagaya Tomoyuki +4
Fair face recognition is all about learning invariant feature that generalizes to unseen faces in any demographic group. Unfortunately, face datasets inevitably capture the imbalan…
Towards Debiasing Frame Length Bias in Text-Video Retrieval via Causal Intervention
Burak Satar, Hongyuan Zhu, Hanwang Zhang +1
Many studies focus on improving pretraining or developing new backbones in text-video retrieval. However, existing methods may suffer from the learning and inference bias issue, as…
An Overview of Challenges in Egocentric Text-Video Retrieval
Burak Satar, Hongyuan Zhu, Hanwang Zhang +1
Text-video retrieval contains various challenges, including biases coming from diverse sources. We highlight some of them supported by illustrations to open a discussion. Besides,…
Unbiased Multiple Instance Learning for Weakly Supervised Video Anomaly Detection
Hui Lv, Zhongqi Yue, Qianru Sun +3
Weakly Supervised Video Anomaly Detection (WSVAD) is challenging because the binary anomaly label is only given on the video level, but the output requires snippet-level prediction…