7 papers
Autoregressive Visual Generation Needs a Prologue
Bowen Zheng, Weijian Luo, Guang Yang +2
In this work, we propose Prologue, an approach to bridging the reconstruction-generation gap in autoregressive (AR) image generation. Instead of modifying visual tokens to satisfy…
Taming the Entropy Cliff: Variable Codebook Size Quantization for Autoregressive Visual Generation
Bowen Zheng, Weijian Luo, Guang Yang +2
Most discrete visual tokenizers rely on a default design: every position in the sequence shares the same codebook. Researchers try to scale the codebook size to get better reco…
Coupling AI and Citizen Science in Creation of Enhanced Training Dataset for Medical Image Segmentation
Amir Syahmi, Xiangrong Lu, Yinxuan Li +7
Recent advancements in medical imaging and artificial intelligence (AI) have greatly enhanced diagnostic capabilities, but the development of effective deep learning (DL) models is…
RSFR: A Coarse-to-Fine Reconstruction Framework for Diffusion Tensor Cardiac MRI with Semantic-Aware Refinement
Jiahao Huang, Fanwen Wang, Pedro F. Ferreira +14
Cardiac diffusion tensor imaging (DTI) offers unique insights into cardiomyocyte arrangements, bridging the gap between microscopic and macroscopic cardiac function. However, its c…
Beyond the Hype: A dispassionate look at vision-language models in medical scenario
Yang Nan, Huichi Zhou, Xiaodan Xing +1
Recent advancements in Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities across diverse tasks, garnering significant attention in AI communities. Howev…
Deep Generative Models Unveil Patterns in Medical Images Through Vision-Language Conditioning
Xiaodan Xing, Junzhi Ning, Yang Nan +1
Deep generative models have significantly advanced medical imaging analysis by enhancing dataset size and quality. Beyond mere data augmentation, our research in this paper highlig…