6 papers
MicroWorld: Empowering Multimodal Large Language Models to Bridge the Microscopic Domain Gap with Multimodal Attribute Graph
Manyu Li, Ruian He, Chenxi Ma +2
Multimodal large language models (MLLMs) show remarkable potential for scientific reasoning, yet their performance in specialized domains such as microscopy remains limited by the…
Non-Uniform Class-Wise Coreset Selection for Vision Model Fine-tuning
Hanyu Zhang, Zhen Xing, Ruian He +4
Coreset selection aims to identify a small yet highly informative subset of data, thereby enabling more efficient model training while reducing storage overhead. Recently, this cap…
MicroVQA++: High-Quality Microscopy Reasoning Dataset with Weakly Supervised Graphs for Multimodal Large Language Model
Manyu Li, Ruian He, Chenxi Ma +2
Multimodal Large Language Models are increasingly applied to biomedical imaging, yet scientific reasoning for microscopy remains limited by the scarcity of large-scale, high-qualit…
Unifying Segment Anything in Microscopy with Vision-Language Knowledge
Manyu Li, Ruian He, Zixian Zhang +3
Accurate segmentation of regions of interest in biomedical images holds substantial value in image analysis. Although several foundation models for biomedical segmentation have cur…
MM-Skin: Enhancing Dermatology Vision-Language Model with an Image-Text Dataset Derived from Textbooks
Wenqi Zeng, Yuqi Sun, Chenxi Ma +2
Medical vision-language models (VLMs) have shown promise as clinical assistants across various medical fields. However, specialized dermatology VLM capable of delivering profession…
Scaling Laws for Data-Efficient Visual Transfer Learning
Wenxuan Yang, Qingqu Wei, Chenxi Ma +2
Current scaling laws for visual AI models focus predominantly on large-scale pretraining, leaving a critical gap in understanding how performance scales for data-constrained downst…