1 citations · 1 across the 4 of their papers we have counts for
10 papers
Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation
Jiawei Mao, Yuhan Wang, Lifeng Chen +6
Recent advances in generative medical models are constrained by modality-specific scenarios that hinder the integration of complementary evidence from imaging, pathology, and clini…
DSKC: Domain Style Modeling with Adaptive Knowledge Consolidation for Exemplar-free Lifelong Person Re-Identification
Shiben Liu, Mingyue Xu, Huijie Fan +3
Lifelong Person Re-identification (LReID) aims to continuously match individuals across camera views from sequential data streams. Existing LReID methods often ignore domain-specif…
ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales
Yihao Zhen, Qiang Wang, Yu Qiao +2
A main challenge of Visual-Language Tracking (VLT) is the misalignment between visual inputs and language descriptions caused by target movement. Previous trackers have explored ma…
FedVLMBench: Benchmarking Federated Fine-Tuning of Vision-Language Models
Weiying Zheng, Ziyue Lin, Pengxin Guo +3
Vision-Language Models (VLMs) have demonstrated remarkable capabilities in cross-modal understanding and generation by integrating visual and textual information. While instruction…
Distribution-aware Forgetting Compensation for Exemplar-Free Lifelong Person Re-identification
Shiben Liu, Huijie Fan, Qiang Wang +3
Lifelong Person Re-identification (LReID) suffers from a key challenge in preserving old knowledge while adapting to new information. The existing solutions include rehearsal-based…
DVG-Diffusion: Dual-View Guided Diffusion Model for CT Reconstruction from X-Rays
Xing Xie, Jiawei Liu, Huijie Fan +3
Directly reconstructing 3D CT volume from few-view 2D X-rays using an end-to-end deep learning network is a challenging task, as X-ray images are merely projection views of the 3D…