6 papers
OmniFashion: Towards Generalist Fashion Intelligence via Multi-Task Vision-Language Learning
Zhengwei Yang, Andi Long, Hao Li +3
Fashion intelligence spans multiple tasks, i.e., retrieval, recommendation, recognition, and dialogue, yet remains hindered by fragmented supervision and incomplete fashion annotat…
Expanding Zero-Shot Object Counting with Rich Prompts
Huilin Zhu, Senyao Li, Jingling Yuan +5
Expanding pre-trained zero-shot counting models to handle unseen categories requires more than simply adding new prompts, as this approach does not achieve the necessary alignment…
VEGAS: Towards Visually Explainable and Grounded Artificial Social Intelligence
Hao Li, Hao Fei, Zechao Hu +2
Social Intelligence Queries (Social-IQ) serve as the primary multimodal benchmark for evaluating a model's social intelligence level. While impressive multiple-choice question(MCQ)…
Cross-Modal Few-Shot Learning: a Generative Transfer Learning Framework
Zhengwei Yang, Yuke Li, Qiang Sun +3
Most existing studies on few-shot learning focus on unimodal settings, where models are trained to generalize to unseen data using a limited amount of labeled examples from a singl…
FocalCount: Towards Class-Count Imbalance in Class-Agnostic Counting
Huilin Zhu, Jingling Yuan, Zhengwei Yang +3
In class-agnostic object counting, the goal is to estimate the total number of object instances in an image without distinguishing between specific categories. Existing methods oft…
Zero-shot Object Counting with Good Exemplars
Huilin Zhu, Jingling Yuan, Zhengwei Yang +4
Zero-shot object counting (ZOC) aims to enumerate objects in images using only the names of object classes during testing, without the need for manual annotations. However, a criti…