activity
20242026
most citedSee Detail Say Clear: Towards Brain CT Report Generation via Pathological Clue-driven Representation Learning

1 citations · 1 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV2026

Detached Skip-Links and -Probe: Decoupling Feature Aggregation from Gradient Propagation for MLLM OCR

Ziye Yuan, Ruchang Yao, Chengxin Zheng +3

Multimodal large language models (MLLMs) excel at high-level reasoning yet fail on OCR tasks where fine-grained visual details are compromised or misaligned. We identify an overloo…

cs.CV2025

Qianfan-VL: Domain-Enhanced Universal Vision-Language Models

Daxiang Dong, Mingming Zheng, Dong Xu +32

We present Qianfan-VL, a series of multimodal large language models ranging from 3B to 70B parameters, achieving state-of-the-art performance through innovative domain enhancement…

cs.CV2025

HSENet: Hybrid Spatial Encoding Network for 3D Medical Vision-Language Understanding

Yanzhao Shi, Xiaodan Zhang, Junzhong Ji +4

Automated 3D CT diagnosis empowers clinicians to make timely, evidence-based decisions by enhancing diagnostic accuracy and workflow efficiency. While multimodal large language mod…

cs.AI2025

MEPNet: Medical Entity-balanced Prompting Network for Brain CT Report Generation

Xiaodan Zhang, Yanzhao Shi, Junzhong Ji +2

The automatic generation of brain CT reports has gained widespread attention, given its potential to assist radiologists in diagnosing cranial diseases. However, brain CT scans inv…

cs.CV20241 cited

See Detail Say Clear: Towards Brain CT Report Generation via Pathological Clue-driven Representation Learning

Chengxin Zheng, Junzhong Ji, Yanzhao Shi +2

Brain CT report generation is significant to aid physicians in diagnosing cranial diseases. Recent studies concentrate on handling the consistency between visual and textual pathol…