papers

Publications (19)

cs.AI2026

Can MLLMs Read Students' Minds? Unpacking Multimodal Error Analysis in Handwritten Math

Dingjie Song, Tianlong Xu, Yi-Fan Zhang +6

Assessing student handwritten scratchwork is crucial for personalized educational feedback but presents unique challenges due to diverse handwriting, complex layouts, and varied pr…

cs.CV2023

MA-SAM: Modality-agnostic SAM Adaptation for 3D Medical Image Segmentation

Cheng Chen, Juzheng Miao, Dufan Wu +10

The Segment Anything Model (SAM), a foundation model for general image segmentation, has demonstrated impressive zero-shot performance across numerous natural image segmentation ta…

cs.AI2026

Agents' Last Exam

Yiyou Sun, Xinyang Han, Weichen Zhang +306

Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…

cs.CV2025

SAMed-2: Selective Memory Enhanced Medical Segment Anything Model

Zhiling Yan, Sifan Song, Dingjie Song +11

Recent "segment anything" efforts show promise by learning from large-scale data, but adapting such models directly to medical images remains challenging due to the complexity of m…

eess.IV2022

Shuffle Instances-based Vision Transformer for Pancreatic Cancer ROSE Image Classification

Tianyi Zhang, Youdan Feng, Yunlu Feng +6

The rapid on-site evaluation (ROSE) technique can signifi-cantly accelerate the diagnosis of pancreatic cancer by im-mediately analyzing the fast-stained cytopathological images. C…

cs.AI2026

LiveMedBench: A Contamination-Free Medical Benchmark for LLMs with Automated Rubric Evaluation

Zhiling Yan, Dingjie Song, Zhe Fang +4

The deployment of Large Language Models (LLMs) in high-stakes clinical settings demands rigorous and reliable evaluation. However, existing medical benchmarks remain static, suffer…

cs.CV2024

SAMAug: Point Prompt Augmentation for Segment Anything Model

Haixing Dai, Chong Ma, Zhiling Yan +14

This paper introduces SAMAug, a novel visual point augmentation method for the Segment Anything Model (SAM) that enhances interactive image segmentation performance. SAMAug generat…

cs.CV2023

Learning to Generate Poetic Chinese Landscape Painting with Calligraphy

Shaozu Yuan, Aijun Dai, Zhiling Yan +5

In this paper, we present a novel system (denoted as Polaca) to generate poetic Chinese landscape painting with calligraphy. Unlike previous single image-to-image painting generati…

cs.AI2026

OpenSkill: Open-World Self-Evolution for LLM Agents

Zhiling Yan, Dingjie Song, Hanrong Zhang +8

Self-evolving agents requires adaptation after deployment, but existing approaches assume a usable learning loop, such as curated skills, successful trajectories, or verifier signa…

cs.CV2026

SegMix:Shuffle-based Feedback Learning for Semantic Segmentation of Pathology Images

Zhiling Yan, Sicheng Chen, Tianyi Zhang +3

Segmentation is a critical task in computational pathology, as it identifies areas affected by disease or abnormal growth and is essential for diagnosis and treatment. However, acq…

cs.CY2026

Agentic AI Enhances Physician Trust in Clinical Decision Making

Zhiling Yan, Zhe Fang, David J King +10

Medical AI has shifted from reasoning to agentic AI, a new paradigm that autonomously invokes external tools during reasoning, rendering intermediate reasoning steps and tool outpu…

eess.IV2024

Medical Unlearnable Examples: Securing Medical Data from Unauthorized Training via Sparsity-Aware Local Masking

Weixiang Sun, Yixin Liu, Zhiling Yan +2

The rapid expansion of AI in healthcare has led to a surge in medical data generation and storage, boosting medical AI development. However, fears of unauthorized use, like trainin…

cs.CV2023

Multimodal ChatGPT for Medical Applications: an Experimental Study of GPT-4V

Zhiling Yan, Kai Zhang, Rong Zhou +3

In this paper, we critically evaluate the capabilities of the state-of-the-art multimodal large language model, i.e., GPT-4 with Vision (GPT-4V), on Visual Question Answering (VQA)…

cs.CV2023

CellMix: A General Instance Relationship based Method for Data Augmentation Towards Pathology Image Classification

Tianyi Zhang, Zhiling Yan, Chunhui Li +5

In pathology image analysis, obtaining and maintaining high-quality annotated samples is an extremely labor-intensive task. To overcome this challenge, mixing-based methods have em…

eess.IV2024

TTT-Unet: Enhancing U-Net with Test-Time Training Layers for Biomedical Image Segmentation

Rong Zhou, Zhengqing Yuan, Zhiling Yan +7

Biomedical image segmentation is crucial for accurately diagnosing and analyzing various diseases. However, Convolutional Neural Networks (CNNs) and Transformers, the most commonly…

cs.CV2024

Biomedical SAM 2: Segment Anything in Biomedical Images and Videos

Zhiling Yan, Weixiang Sun, Rong Zhou +8

Medical image segmentation and video object segmentation are essential for diagnosing and analyzing diseases by identifying and measuring biological structures. Recent advances in…

cs.CV2024

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Yixin Liu, Kai Zhang, Yuan Li +9

Sora is a text-to-video generative AI model, released by OpenAI in February 2024. The model is trained to generate videos of realistic or imaginative scenes from text instructions…

cs.CL2024

BiomedGPT: A Generalist Vision-Language Foundation Model for Diverse Biomedical Tasks

Kai Zhang, Rong Zhou, Eashan Adhikarla +20

Traditional biomedical artificial intelligence (AI) models, designed for specific tasks or modalities, often exhibit limited flexibility in real-world deployment and struggle to ut…

cs.AI2023

A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT

Yihan Cao, Siyu Li, Yixin Liu +4

Recently, ChatGPT, along with DALL-E-2 and Codex,has been gaining significant attention from society. As a result, many individuals have become interested in related resources and…