10 papers
Loom: Diffusion-Transformer for Interleaved Generation
Mingcheng Ye, Jiaming Liu, Yiren Song
Interleaved text-image generation aims to jointly produce coherent visual frames and aligned textual descriptions within a single sequence, enabling tasks such as style transfer, c…
SpinalSAM-R1: A Vision-Language Multimodal Interactive System for Spine CT Segmentation
Jiaming Liu, Dingwei Fan, Junyong Zhao +3
The anatomical structure segmentation of the spine and adjacent structures from computed tomography (CT) images is a key step for spinal disease diagnosis and treatment. However, t…
MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
Liujian Tang, Shaokang Dong, Yijia Huang +21
This paper presents MagicGUI, a foundational mobile GUI agent designed to address critical challenges in perception, grounding, and reasoning within real-world mobile GUI environme…
Awesome-OL: An Extensible Toolkit for Online Learning
Zeyi Liu, Songqiao Hu, Pengyu Han +2
In recent years, online learning has attracted increasing attention due to its adaptive capability to process streaming and non-stationary data. To facilitate algorithm development…
Stable-Hair v2: Real-World Hair Transfer via Multiple-View Diffusion Model
Kuiyuan Sun, Yuxuan Zhang, Jichao Zhang +4
While diffusion-based methods have shown impressive capabilities in capturing diverse and complex hairstyles, their ability to generate consistent and high-quality multi-view outpu…
WordCon: Word-level Typography Control in Scene Text Rendering
Wenda Shi, Yiren Song, Zihan Rao +3
Achieving precise word-level typography control within generated images remains a persistent challenge. To address it, we newly construct a word-level controlled scene text dataset…