54 citations · 63 across the 4 of their papers we have counts for
4 papers
FaceInsight: A Multimodal Large Language Model for Face Perception
Jingzhi Li, Changjiang Luo, Ruoyu Chen +4
Recent advances in multimodal large language models (MLLMs) have demonstrated strong capabilities in understanding general visual content. However, these general-domain MLLMs perfo…
Beyond Progress Measures: Theoretical Insights into the Mechanism of Grokking
Zihan Gu, Ruoyu Chen, Hua Zhang +2
Grokking, referring to the abrupt improvement in test accuracy after extended overfitting, offers valuable insights into the mechanisms of model generalization. Existing researches…
Generalized Semantic Contrastive Learning via Embedding Side Information for Few-Shot Object Detection
Ruoyu Chen, Hua Zhang, Jingzhi Li +3
The objective of few-shot object detection (FSOD) is to detect novel objects with few training samples. The core challenge of this task is how to construct a generalized feature sp…
Lightweight Pixel Difference Networks for Efficient Visual Representation Learning
Zhuo Su, Jiehua Zhang, Longguang Wang +4
Recently, there have been tremendous efforts in developing lightweight Deep Neural Networks (DNNs) with satisfactory accuracy, which can enable the ubiquitous deployment of DNNs in…