collaborators

6 papers

cs.CV2026

What Does Your Short-Answer VQA Score Actually Measure? Evaluator-Dependent Instability in Multimodal Short-Answer Benchmarks

Guanhua Ye, Niu Jingbin, Yan Li +4

Short-answer VQA benchmarks conflate two distinct quantities: whether a model's answer is semantically correct, and whether that answer matches the surface form expected by the aut…

cs.IR2026

Research Team Identification Based on Representation Learning of Academic Heterogeneous Information Network

Junfu Wang, Yawen Li, Zhe Xue +1

Academic networks in the real world can usually be described by heterogeneous information networks composed of multiple types of nodes and relationships. Existing representation-le…

eess.IV2026

Beyond Metadata: CAPRA for Hidden Subgroup Analysis under Missing Metadata in Medical Imaging

Yawen Li, Yan Li, Zhe Xue +3

Medical imaging models are often deployed without the demographic, acquisition, and quality metadata needed for subgroup auditing. Once those metadata disappear, clinically critica…

cs.CL2026

Entity Alignment Method of Science and Technology Patent based on Graph Convolution Network and Information Fusion

Runze Fang, Yawen Li, Yingxia Shao +2

The entity alignment of science and technology patents aims to link the equivalent entities in the knowledge graph of different science and technology patent data sources. Most ent…

cs.CL2026

Semantic Representation Learning of Scientific Literature based on Adaptive Feature and Graph Neural Network

Hongrui Gao, Yawen Li, Meiyu Liang +2

Because most scientific literature data are unlabeled, semantic representation learning based on unsupervised graphs has become crucial. To enrich scientific-literature features, t…

cs.CV2024

Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning

Zhengyang Liang, Meiyu Liang, Wei Huang +2

In recent years, pre-trained multimodal large models have attracted widespread attention due to their outstanding performance in various multimodal applications. Nonetheless, the e…