3 citations · 4 across the 3 of their papers we have counts for
3 papers
cs.CV2026
VisionPangu: A Compact and Fine-Grained Multimodal Assistant with 1.7B Parameters
Jiaxin Fan, Wenpo Song
Large Multimodal Models (LMMs) have achieved strong performance in vision-language understanding, yet many existing approaches rely on large-scale architectures and coarse supervis…
cs.CL2023★ 1 cited
Boosting Chinese ASR Error Correction with Dynamic Error Scaling Mechanism
Jiaxin Fan, Yong Zhang, Hanzhang Li +5
Chinese Automatic Speech Recognition (ASR) error correction presents significant challenges due to the Chinese language's unique features, including a large character set and borde…
eess.IV2022★ 3 cited
PVT-COV19D: Pyramid Vision Transformer for COVID-19 Diagnosis
Lilang Zheng, Jiaxuan Fang, Xiaorun Tang +5
With the outbreak of COVID-19, a large number of relevant studies have emerged in recent years. We propose an automatic COVID-19 diagnosis framework based on lung CT scan images, t…