1 citations · 2 across the 13 of their papers we have counts for
19 papers
Chain of World: World Model Thinking in Latent Motion
Fuxiang Yang, Donglin Di, Lulu Tang +6
Vision-Language-Action (VLA) models are a promising path toward embodied intelligence, yet they often overlook the predictive and temporal-causal structure underlying visual dynami…
Visual Prompt-Agnostic Evolution
Junze Wang, Lei Fan, Dezheng Zhang +5
Visual Prompt Tuning (VPT) adapts a frozen Vision Transformer (ViT) to downstream tasks by inserting a small number of learnable prompt tokens into the token sequence at each layer…
EFDiT: Efficient Fine-grained Image Generation Using Diffusion Transformer Models
Kun Wang, Donglin Di, Tonghua Su +1
Diffusion models are highly regarded for their controllability and the diversity of images they generate. However, class-conditional generation methods based on diffusion models of…
Cross-Stain Contrastive Learning for Paired Immunohistochemistry and Histopathology Slide Representation Learning
Yizhi Zhang, Lei Fan, Zhulin Tao +4
Universal, transferable whole-slide image (WSI) representations are central to computational pathology. Incorporating multiple markers (e.g., immunohistochemistry, IHC) alongside H…
Global-Local Aware Scene Text Editing
Fuxiang Yang, Tonghua Su, Donglin Di +4
Scene Text Editing (STE) involves replacing text in a scene image with new target text while preserving both the original text style and background texture. Existing methods suffer…
GAOT: Generating Articulated Objects Through Text-Guided Diffusion Models
Hao Sun, Lei Fan, Donglin Di +1
Articulated object generation has seen increasing advancements, yet existing models often lack the ability to be conditioned on text prompts. To address the significant gap between…