collaborators

6 papers

cs.CV2026

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

Hangjie Yuan, Yichen Qian, Zhiwei Tang +21

Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric chall…

cs.CV2026

Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models

Jiazheng Xing, Hangjie Yuan, Lingling Cai +9

Connector-based video unified models have demonstrated strong capability in instruction-grounded video synthesis, but integrating a large high-fidelity generator into the unified t…

cs.CV2026

Towards Error-Free Long Video Generation

Shuning Chang, Weihua Chen, Jiasheng Tang +8

Recent advances in video generation have made minute-level synthesis possible; however, generating long videos remains challenging due to error accumulation, attribute drift, and t…

cs.CV2026

Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model Perspective

Hangjie Yuan, Weihua Chen, Jun Cen +11

Autoregressive large language models (LLMs) have unified a vast range of language tasks, inspiring preliminary efforts in autoregressive (AR) video generation. Existing AR video ge…

cs.CV2025

SparseDiT: Token Sparsification for Efficient Diffusion Transformer

Shuning Chang, Pichao Wang, Jiasheng Tang +2

Diffusion Transformers (DiT) are renowned for their impressive generative performance; however, they are significantly constrained by considerable computational costs due to the qu…

cs.LG2025

Flow Along the K-Amplitude for Generative Modeling

Weitao Du, Shuning Chang, Jiasheng Tang +3

In this work, we propose a novel generative learning paradigm, K-Flow, an algorithm that flows along the -amplitude. Here, is a scaling parameter that organizes frequency ba…