collaborators

7 papers

cs.CV2026

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models

Zhiwei Yang, Yuanchen Wu, Nan Zhang +3

The paper proposes Scene Graph Thinking (SaGe), a method that equips multimodal large language models with explicit scene‑graph representations to enable fine‑grained, structured v…

cs.CV2026

DiCLIP: Diffusion Model Enhances CLIP's Dense Knowledge for Weakly Supervised Semantic Segmentation

Zhiwei Yang, Pengfei Song, Yucong Meng +3

Weakly Supervised Semantic Segmentation (WSSS) with image-level labels typically leverages Class Activation Maps (CAMs) to achieve pixel-level predictions. Recently, Contrastive La…

eess.IV2025

DH-Mamba: Exploring Dual-domain Hierarchical State Space Models for MRI Reconstruction

Yucong Meng, Zhiwei Yang, Zhijian Song +1

The accelerated MRI reconstruction poses a challenging ill-posed inverse problem due to the significant undersampling in k-space. Deep neural networks, such as CNNs and ViTs, have…

cs.CV2025

Exploring CLIP's Dense Knowledge for Weakly Supervised Semantic Segmentation

Zhiwei Yang, Yucong Meng, Kexue Fu +3

Weakly Supervised Semantic Segmentation (WSSS) with image-level labels aims to achieve pixel-level predictions using Class Activation Maps (CAMs). Recently, Contrastive Language-Im…

eess.IV2025

Continuous K-space Recovery Network with Image Guidance for Fast MRI Reconstruction

Yucong Meng, Zhiwei Yang, Minghong Duan +2

Magnetic resonance imaging (MRI) is a crucial tool for clinical diagnosis while facing the challenge of long scanning time. To reduce the acquisition time, fast MRI reconstruction…

cs.CV2025

MoRe: Class Patch Attention Needs Regularization for Weakly Supervised Semantic Segmentation

Zhiwei Yang, Yucong Meng, Kexue Fu +2

Weakly Supervised Semantic Segmentation (WSSS) with image-level labels typically uses Class Activation Maps (CAM) to achieve dense predictions. Recently, Vision Transformer (ViT) h…