collaborators

8 papers

cs.LG2026

Deep Unfolding: Recent Developments, Theory, and Design Guidelines

Nir Shlezinger, Santiago Segarra, Yi Zhang +4

Optimization methods play a central role in signal processing, serving as the mathematical foundation for inference, estimation, and control. While classical iterative optimization…

cs.CV2026

NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation

Huichao Zhang, Liao Qu, Yiheng Liu +33

We present NextFlow, a unified decoder-only autoregressive transformer trained on 6 trillion interleaved text-image discrete tokens. By leveraging a unified vision representation w…

cs.CV2025

SpatialVID: A Large-Scale Video Dataset with Spatial Annotations

Jiahao Wang, Yufeng Yuan, Rujie Zheng +12

Significant progress has been made in spatial intelligence, spanning both spatial reconstruction and world exploration. However, the scalability and real-world fidelity of current…

cs.SD2025

GLM-TTS Technical Report

Jiayan Cui, Zhihan Yang, Naihan Li +10

This work proposes GLM-TTS, a production-level TTS system designed for efficiency, controllability, and high-fidelity speech generation. GLM-TTS follows a two-stage architecture, c…

cs.CV2025

RAGDiffusion: Faithful Cloth Generation via External Knowledge Assimilation

Xianfeng Tan, Yuhan Li, Wenxiang Shang +6

Standard clothing asset generation involves restoring forward-facing flat-lay garment images displayed on a clear background by extracting clothing information from diverse real-wo…

cs.AI2025

MM-Food-100K: A 100,000-Sample Multimodal Food Intelligence Dataset with Verifiable Provenance

Yi Dong, Yusuke Muraoka, Scott Shi +1

We present MM-Food-100K, a public 100,000-sample multimodal food intelligence dataset with verifiable provenance. It is a curated approximately 10% open subset of an original 1.2 m…