activity
20242026
most citedMemory-Augmented Incomplete Multimodal Survival Prediction via Cross-Slide and Gene-Attentive Hypergraph Learning

1 citations · 2 across the 13 of their papers we have counts for

collaborators

19 papers

cs.CV2026

Chain of World: World Model Thinking in Latent Motion

Fuxiang Yang, Donglin Di, Lulu Tang +6

Vision-Language-Action (VLA) models are a promising path toward embodied intelligence, yet they often overlook the predictive and temporal-causal structure underlying visual dynami…

cs.CV2026

Visual Prompt-Agnostic Evolution

Junze Wang, Lei Fan, Dezheng Zhang +5

Visual Prompt Tuning (VPT) adapts a frozen Vision Transformer (ViT) to downstream tasks by inserting a small number of learnable prompt tokens into the token sequence at each layer…

cs.CV2025

EFDiT: Efficient Fine-grained Image Generation Using Diffusion Transformer Models

Kun Wang, Donglin Di, Tonghua Su +1

Diffusion models are highly regarded for their controllability and the diversity of images they generate. However, class-conditional generation methods based on diffusion models of…

cs.CV2025

Cross-Stain Contrastive Learning for Paired Immunohistochemistry and Histopathology Slide Representation Learning

Yizhi Zhang, Lei Fan, Zhulin Tao +4

Universal, transferable whole-slide image (WSI) representations are central to computational pathology. Incorporating multiple markers (e.g., immunohistochemistry, IHC) alongside H…

cs.CV2025

Global-Local Aware Scene Text Editing

Fuxiang Yang, Tonghua Su, Donglin Di +4

Scene Text Editing (STE) involves replacing text in a scene image with new target text while preserving both the original text style and background texture. Existing methods suffer…

cs.CV2025

GAOT: Generating Articulated Objects Through Text-Guided Diffusion Models

Hao Sun, Lei Fan, Donglin Di +1

Articulated object generation has seen increasing advancements, yet existing models often lack the ability to be conditioned on text prompts. To address the significant gap between…