7 citations · 11 across the 5 of their papers we have counts for
5 papers
VideoCoT: A Video Chain-of-Thought Dataset with Active Annotation Tool
Yan Wang, Yawen Zeng, Jingsheng Zheng +3
Multimodal large language models (MLLMs) are flourishing, but mainly focus on images with less attention than videos, especially in sub-fields such as prompt engineering, video cha…
Energy-based Automated Model Evaluation
Ru Peng, Heming Zou, Haobo Wang +3
The conventional evaluation protocols on machine learning models rely heavily on a labeled, i.i.d-assumed testing dataset, which is not often present in real world applications. Th…
Multi-Prompts Learning with Cross-Modal Alignment for Attribute-based Person Re-Identification
Yajing Zhai, Yawen Zeng, Zhiyong Huang +3
The fine-grained attribute descriptions can significantly supplement the valuable semantic information for person image, which is vital to the success of person re-identification (…
Do LLMs Possess a Personality? Making the MBTI Test an Amazing Evaluation for Large Language Models
Keyu Pan, Yawen Zeng
The field of large language models (LLMs) has made significant progress, and their knowledge storage capacity is approaching that of human beings. Furthermore, advanced techniques,…
Better Sign Language Translation with Monolingual Data
Ru Peng, Yawen Zeng, Junbo Zhao
Sign language translation (SLT) systems, which are often decomposed into video-to-gloss (V2G) recognition and gloss-to-text (G2T) translation through the pivot gloss, heavily relie…