Publications (19)
Can MLLMs Read Students' Minds? Unpacking Multimodal Error Analysis in Handwritten Math
Dingjie Song, Tianlong Xu, Yi-Fan Zhang +6
Assessing student handwritten scratchwork is crucial for personalized educational feedback but presents unique challenges due to diverse handwriting, complex layouts, and varied pr…
MA-SAM: Modality-agnostic SAM Adaptation for 3D Medical Image Segmentation
Cheng Chen, Juzheng Miao, Dufan Wu +10
The Segment Anything Model (SAM), a foundation model for general image segmentation, has demonstrated impressive zero-shot performance across numerous natural image segmentation ta…
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
SAMed-2: Selective Memory Enhanced Medical Segment Anything Model
Zhiling Yan, Sifan Song, Dingjie Song +11
Recent "segment anything" efforts show promise by learning from large-scale data, but adapting such models directly to medical images remains challenging due to the complexity of m…
Shuffle Instances-based Vision Transformer for Pancreatic Cancer ROSE Image Classification
Tianyi Zhang, Youdan Feng, Yunlu Feng +6
The rapid on-site evaluation (ROSE) technique can signifi-cantly accelerate the diagnosis of pancreatic cancer by im-mediately analyzing the fast-stained cytopathological images. C…
LiveMedBench: A Contamination-Free Medical Benchmark for LLMs with Automated Rubric Evaluation
Zhiling Yan, Dingjie Song, Zhe Fang +4
The deployment of Large Language Models (LLMs) in high-stakes clinical settings demands rigorous and reliable evaluation. However, existing medical benchmarks remain static, suffer…
SAMAug: Point Prompt Augmentation for Segment Anything Model
Haixing Dai, Chong Ma, Zhiling Yan +14
This paper introduces SAMAug, a novel visual point augmentation method for the Segment Anything Model (SAM) that enhances interactive image segmentation performance. SAMAug generat…
Learning to Generate Poetic Chinese Landscape Painting with Calligraphy
Shaozu Yuan, Aijun Dai, Zhiling Yan +5
In this paper, we present a novel system (denoted as Polaca) to generate poetic Chinese landscape painting with calligraphy. Unlike previous single image-to-image painting generati…
OpenSkill: Open-World Self-Evolution for LLM Agents
Zhiling Yan, Dingjie Song, Hanrong Zhang +8
Self-evolving agents requires adaptation after deployment, but existing approaches assume a usable learning loop, such as curated skills, successful trajectories, or verifier signa…
SegMix:Shuffle-based Feedback Learning for Semantic Segmentation of Pathology Images
Zhiling Yan, Sicheng Chen, Tianyi Zhang +3
Segmentation is a critical task in computational pathology, as it identifies areas affected by disease or abnormal growth and is essential for diagnosis and treatment. However, acq…
Agentic AI Enhances Physician Trust in Clinical Decision Making
Zhiling Yan, Zhe Fang, David J King +10
Medical AI has shifted from reasoning to agentic AI, a new paradigm that autonomously invokes external tools during reasoning, rendering intermediate reasoning steps and tool outpu…
Medical Unlearnable Examples: Securing Medical Data from Unauthorized Training via Sparsity-Aware Local Masking
Weixiang Sun, Yixin Liu, Zhiling Yan +2
The rapid expansion of AI in healthcare has led to a surge in medical data generation and storage, boosting medical AI development. However, fears of unauthorized use, like trainin…
Multimodal ChatGPT for Medical Applications: an Experimental Study of GPT-4V
Zhiling Yan, Kai Zhang, Rong Zhou +3
In this paper, we critically evaluate the capabilities of the state-of-the-art multimodal large language model, i.e., GPT-4 with Vision (GPT-4V), on Visual Question Answering (VQA)…
CellMix: A General Instance Relationship based Method for Data Augmentation Towards Pathology Image Classification
Tianyi Zhang, Zhiling Yan, Chunhui Li +5
In pathology image analysis, obtaining and maintaining high-quality annotated samples is an extremely labor-intensive task. To overcome this challenge, mixing-based methods have em…
TTT-Unet: Enhancing U-Net with Test-Time Training Layers for Biomedical Image Segmentation
Rong Zhou, Zhengqing Yuan, Zhiling Yan +7
Biomedical image segmentation is crucial for accurately diagnosing and analyzing various diseases. However, Convolutional Neural Networks (CNNs) and Transformers, the most commonly…
Biomedical SAM 2: Segment Anything in Biomedical Images and Videos
Zhiling Yan, Weixiang Sun, Rong Zhou +8
Medical image segmentation and video object segmentation are essential for diagnosing and analyzing diseases by identifying and measuring biological structures. Recent advances in…
Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models
Yixin Liu, Kai Zhang, Yuan Li +9
Sora is a text-to-video generative AI model, released by OpenAI in February 2024. The model is trained to generate videos of realistic or imaginative scenes from text instructions…
BiomedGPT: A Generalist Vision-Language Foundation Model for Diverse Biomedical Tasks
Kai Zhang, Rong Zhou, Eashan Adhikarla +20
Traditional biomedical artificial intelligence (AI) models, designed for specific tasks or modalities, often exhibit limited flexibility in real-world deployment and struggle to ut…
A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT
Yihan Cao, Siyu Li, Yixin Liu +4
Recently, ChatGPT, along with DALL-E-2 and Codex,has been gaining significant attention from society. As a result, many individuals have become interested in related resources and…