6 papers
Towards Generalized Multi-Image Editing for Unified Multimodal Models
Pengcheng Xu, Peng Tang, Donghao Luo +7
Unified Multimodal Models (UMMs) integrate multimodal understanding and generation, yet they are limited to maintaining visual consistency and disambiguating visual cues when refer…
Event-Driven Online Vertical Federated Learning
Ganyu Wang, Boyu Wang, Bin Gu +1
Online learning is more adaptable to real-world scenarios in Vertical Federated Learning (VFL) compared to offline learning. However, integrating online learning into VFL presents…
Textualize Visual Prompt for Image Editing via Diffusion Bridge
Pengcheng Xu, Qingnan Fan, Fei Kou +5
Visual prompt, a pair of before-and-after edited images, can convey indescribable imagery transformations and prosper in image editing. However, current visual prompt methods rely…
Enhancing Generalization in Chain of Thought Reasoning for Smaller Models
Maxwell J. Yin, Dingyi Jiang, Yongbing Chen +2
Chain-of-Thought (CoT) reasoning in smaller language models is a challenging natural language process problem yet highly desirable in many real-life applications. Existing CoT know…
Leveraging Group Classification with Descending Soft Labeling for Deep Imbalanced Regression
Ruizhi Pu, Gezheng Xu, Ruiyi Fang +3
Deep imbalanced regression (DIR), where the target values have a highly skewed distribution and are also continuous, is an intriguing yet under-explored problem in machine learning…
Unveil Inversion and Invariance in Flow Transformer for Versatile Image Editing
Pengcheng Xu, Boyuan Jiang, Xiaobin Hu +7
Leveraging the large generative prior of the flow transformer for tuning-free image editing requires authentic inversion to project the image into the model's domain and a flexible…