most citedGeneralized Semantic Contrastive Learning via Embedding Side Information for Few-Shot Object Detection

9 citations · 9 across the 4 of their papers we have counts for

collaborators

6 papers

cs.CV2025

PhaseWin Search Framework Enable Efficient Object-Level Interpretation

Zihan Gu, Ruoyu Chen, Junchi Zhang +3

Attribution is essential for interpreting object-level foundation models. Recent methods based on submodular subset selection have achieved high faithfulness, but their efficiency…

cs.CV2025

Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding

Youze Wang, Zijun Chen, Ruoyu Chen +8

Recent advancements in multimodal large language models for video understanding (videoLLMs) have enhanced their capacity to process complex spatiotemporal data. However, challenges…

cs.CV2025

FaceInsight: A Multimodal Large Language Model for Face Perception

Jingzhi Li, Changjiang Luo, Ruoyu Chen +4

Recent advances in multimodal large language models (MLLMs) have demonstrated strong capabilities in understanding general visual content. However, these general-domain MLLMs perfo…

cs.LG2025

Beyond Progress Measures: Theoretical Insights into the Mechanism of Grokking

Zihan Gu, Ruoyu Chen, Hua Zhang +2

Grokking, referring to the abrupt improvement in test accuracy after extended overfitting, offers valuable insights into the mechanisms of model generalization. Existing researches…

cs.CV20259 cited

Generalized Semantic Contrastive Learning via Embedding Side Information for Few-Shot Object Detection

Ruoyu Chen, Hua Zhang, Jingzhi Li +3

The objective of few-shot object detection (FSOD) is to detect novel objects with few training samples. The core challenge of this task is how to construct a generalized feature sp…

cs.CV2024

Interpreting Object-level Foundation Models via Visual Precision Search

Ruoyu Chen, Siyuan Liang, Jingzhi Li +5

Advances in multimodal pre-training have propelled object-level foundation models, such as Grounding DINO and Florence-2, in tasks like visual grounding and object detection. Howev…