collaborators

9 papers

cs.CV2026

UR-MLLM: Uncertainty-aware Revisit Reasoning in Multimodal Large Language Models for Radiology Report Generation

Yucheng Chen, Yang Yu, Jiazhou Zhou +5

Radiologists generate diagnostic reports through iterative and selective revisiting of suspicious regions to refine their interpretations. Recent multimodal large language models (…

cs.CV2026

Seeing Through Multiple Views: Parameter-Efficient Fine-Tuning via Selective Neurons for Consistent Radiology Report Generation

Yucheng Chen, Jinjing Zhu, Yang Yu +7

Recent years have seen substantial advances in radiology report generation (RRG), yet existing approaches predominantly adopt direct feature fusion when handling multi-view X-ray i…

cs.CV2026

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images

Qirui Wang, Jingyi He, Yining Pan +3

Spatial reasoning (SR), the ability to infer 3D spatial information from 2D inputs, is essential for real-world applications such as embodied AI and autonomous driving. However, ex…

cs.CV2026

One Sentence, One Drama: Personalized Short-Form Drama Generation via Multi-Agent Systems

Yufei Shi, Weilong Yan, Naixuan Huang +5

Existing approaches for digital short-drama production typically rely on one-shot LLM generated scripts and loosely coupled pipelines, which fail to satisfy three key requirements…

cs.CV2026

RIHA: Report-Image Hierarchical Alignment for Radiology Report Generation

Yucheng Chen, Yang Yu, Yufei Shi +3

Radiology report generation (RRG) has emerged as a promising approach to alleviate radiologists' workload and reduce human errors by automatically generating diagnostic reports fro…

cs.CV2026

Exploiting Low-Dimensional Manifold of Features for Few-Shot Whole Slide Image Classification

Conghao Xiong, Zhengrui Guo, Zhe Xu +6

Few-shot Whole Slide Image (WSI) classification is severely hampered by overfitting. We argue that this is not merely a data-scarcity issue but a fundamentally geometric problem. G…