13 papers
InnoText: A Unified Model for Visual Text Generation and Editing
Haowei Liu, Runze He, Jian Lu +10
Diffusion models have recently achieved remarkable success in high-fidelity image synthesis, yet their application to visual text generation and editing remains relatively underexp…
Keep It in Mind: User Centric Continual Spatial Intelligence Reasoning in Egocentric Video Streams
Yun Wang, Junbin Xiao, Han Lyu +6
We introduce UCS-Bench, a dataset spanning 170+ hours of egocentric visual observations with 8.1K+ timestamped questions for diagnosing User-Centric Continual Spatial intelligence…
FedCVU: Federated Learning for Cross-View Video Understanding
Shenghan Zhang, Run Ling, Ke Cao +2
Federated learning (FL) has emerged as a promising paradigm for privacy-preserving multi-camera video understanding. However, applying FL to cross-view scenarios faces three major…
InnoAds-Composer: Efficient Condition Composition for E-Commerce Poster Generation
Yuxin Qin, Ke Cao, Haowei Liu +13
E-commerce product poster generation aims to automatically synthesize a single image that effectively conveys product information by presenting a subject, text, and a designed styl…
Cross-Scale Pansharpening via ScaleFormer and the PanScale Benchmark
Ke Cao, Xuanhua He, Xueheng Li +7
Pansharpening aims to generate high-resolution multi-spectral images by fusing the spatial detail of panchromatic images with the spectral richness of low-resolution MS data. Howev…
StyMam: A Mamba-Based Generator for Artistic Style Transfer
Zhou Hong, Ning Dong, Yicheng Di +8
Image style transfer aims to integrate the visual patterns of a specific artistic style into a content image while preserving its content structure. Existing methods mainly rely on…