9 citations · 9 across the 4 of their papers we have counts for
6 papers
PhaseWin Search Framework Enable Efficient Object-Level Interpretation
Zihan Gu, Ruoyu Chen, Junchi Zhang +3
Attribution is essential for interpreting object-level foundation models. Recent methods based on submodular subset selection have achieved high faithfulness, but their efficiency…
Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding
Youze Wang, Zijun Chen, Ruoyu Chen +8
Recent advancements in multimodal large language models for video understanding (videoLLMs) have enhanced their capacity to process complex spatiotemporal data. However, challenges…
FaceInsight: A Multimodal Large Language Model for Face Perception
Jingzhi Li, Changjiang Luo, Ruoyu Chen +4
Recent advances in multimodal large language models (MLLMs) have demonstrated strong capabilities in understanding general visual content. However, these general-domain MLLMs perfo…
Beyond Progress Measures: Theoretical Insights into the Mechanism of Grokking
Zihan Gu, Ruoyu Chen, Hua Zhang +2
Grokking, referring to the abrupt improvement in test accuracy after extended overfitting, offers valuable insights into the mechanisms of model generalization. Existing researches…
Generalized Semantic Contrastive Learning via Embedding Side Information for Few-Shot Object Detection
Ruoyu Chen, Hua Zhang, Jingzhi Li +3
The objective of few-shot object detection (FSOD) is to detect novel objects with few training samples. The core challenge of this task is how to construct a generalized feature sp…
Interpreting Object-level Foundation Models via Visual Precision Search
Ruoyu Chen, Siyuan Liang, Jingzhi Li +5
Advances in multimodal pre-training have propelled object-level foundation models, such as Grounding DINO and Florence-2, in tasks like visual grounding and object detection. Howev…