9 papers · 1 filter
Loom: Diffusion-Transformer for Interleaved Generation
Mingcheng Ye, Jiaming Liu, Yiren Song
Interleaved text-image generation aims to jointly produce coherent visual frames and aligned textual descriptions within a single sequence, enabling tasks such as style transfer, c…
SpinalSAM-R1: A Vision-Language Multimodal Interactive System for Spine CT Segmentation
Jiaming Liu, Dingwei Fan, Junyong Zhao +3
The anatomical structure segmentation of the spine and adjacent structures from computed tomography (CT) images is a key step for spinal disease diagnosis and treatment. However, t…
MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
Liujian Tang, Shaokang Dong, Yijia Huang +21
This paper presents MagicGUI, a foundational mobile GUI agent designed to address critical challenges in perception, grounding, and reasoning within real-world mobile GUI environme…
Awesome-OL: An Extensible Toolkit for Online Learning
Zeyi Liu, Songqiao Hu, Pengyu Han +2
In recent years, online learning has attracted increasing attention due to its adaptive capability to process streaming and non-stationary data. To facilitate algorithm development…
FonTS: Text Rendering with Typography and Style Controls
Wenda Shi, Yiren Song, Dengming Zhang +2
Visual text rendering are widespread in various real-world applications, requiring careful font selection and typographic choices. Recent progress in diffusion transformer (DiT)-ba…
Stable-Hair v2: Real-World Hair Transfer via Multiple-View Diffusion Model
Kuiyuan Sun, Yuxuan Zhang, Jichao Zhang +4
While diffusion-based methods have shown impressive capabilities in capturing diverse and complex hairstyles, their ability to generate consistent and high-quality multi-view outpu…