11 papers
Exploring Cross-Modal Flows for Few-Shot Learning
Ziqi Jiang, Yanghao Wang, Long Chen
Aligning features from different modalities, is one of the most fundamental challenges for cross-modal tasks. Although pre-trained vision-language models can achieve a general alig…
MotionAdapter: Video Motion Transfer via Content-Aware Attention Customization
Zhexin Zhang, Yangyang Xu, Yifeng Zhu +4
Recent advances in diffusion-based text-to-video models, particularly those built on the diffusion transformer architecture, have achieved remarkable progress in generating high-qu…
Target-aware Image Editing via Cycle-consistent Constraints
Yanghao Wang, Zhen Wang, Long Chen
Recent pre-trained text-to-image flow models have enabled remarkable progress in text-based image editing. Mainstream approaches adopt a corruption-then-restoration paradigm, where…
Adversarial Batch Representation Augmentation for Batch Correction in High-Content Cellular Screening
Lei Tong, Xujing Yao, Adam Corrigan +6
High-Content Screening routinely generates massive volumes of cell painting images for phenotypic profiling. However, technical variations across experimental executions inevitably…
LLM-Guided Diagnostic Evidence Alignment for Medical Vision-Language Pretraining under Limited Pairing
Huimin Yan, Liang Bai, Xian Yang +1
Most existing CLIP-style medical vision--language pretraining methods rely on global or local alignment with substantial paired data. However, global alignment is easily dominated…
FlowDC: Flow-Based Decoupling-Decay for Complex Image Editing
Yilei Jiang, Zhen Wang, Yanghao Wang +4
With the surge of pre-trained text-to-image flow matching models, text-based image editing performance has gained remarkable improvement, especially for \underline{simple editing}…