1 citations · 2 across the 10 of their papers we have counts for
6 papers · 1 filter
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm
Yaofang Liu, Kangning Cui, Meng Chu +7
Humans often specify and create through visual artifacts: typography sheets, sketches, reference images, and annotated scenes. Yet modern visual generators still ask users to seria…
Learning Where to Embed: Noise-Aware Positional Embedding for Query Retrieval in Small-Object Detection
Yangchen Zeng, Zhenyu Yu, Dongming Jiang +5
Transformer-based detectors have advanced small-object detection, but they often remain inefficient and vulnerable to background-induced query noise, which motivates deep decoders…
StoryState: Agent-Based State Control for Consistent and Editable Storybooks
Ayushman Sarkar, Zhenyu Yu, Wei Tang +3
Large multimodal models have enabled one-click storybook generation, where users provide a short description and receive a multi-page illustrated story. However, the underlying sto…
ReDiStory: Region-Disentangled Diffusion for Consistent Visual Story Generation
Ayushman Sarkar, Zhenyu Yu, Chu Chen +3
Generating coherent visual stories requires maintaining subject identity across multiple images while preserving frame-specific semantics. Recent training-free methods concatenate…
MoSAiC: Multi-Modal Multi-Label Supervision-Aware Contrastive Learning for Remote Sensing
Debashis Gupta, Aditi Golder, Rongkhun Zhu +6
Contrastive learning (CL) has emerged as a powerful paradigm for learning transferable representations without the reliance on large labeled datasets. Its ability to capture intrin…
Center-guided Classifier for Semantic Segmentation of Remote Sensing Images
Wei Zhang, Mengting Ma, Yizhen Jiang +4
Compared with natural images, remote sensing images (RSIs) have the unique characteristic. i.e., larger intraclass variance, which makes semantic segmentation for remote sensing im…