2 citations · 2 across the 5 of their papers we have counts for
8 papers
NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
Huichao Zhang, Liao Qu, Yiheng Liu +33
We present NextFlow, a unified decoder-only autoregressive transformer trained on 6 trillion interleaved text-image discrete tokens. By leveraging a unified vision representation w…
GLM-TTS Technical Report
Jiayan Cui, Zhihan Yang, Naihan Li +10
This work proposes GLM-TTS, a production-level TTS system designed for efficiency, controllability, and high-fidelity speech generation. GLM-TTS follows a two-stage architecture, c…
SpatialVID: A Large-Scale Video Dataset with Spatial Annotations
Jiahao Wang, Yufeng Yuan, Rujie Zheng +12
Significant progress has been made in spatial intelligence, spanning both spatial reconstruction and world exploration. However, the scalability and real-world fidelity of current…
MM-Food-100K: A 100,000-Sample Multimodal Food Intelligence Dataset with Verifiable Provenance
Yi Dong, Yusuke Muraoka, Scott Shi +1
We present MM-Food-100K, a public 100,000-sample multimodal food intelligence dataset with verifiable provenance. It is a curated approximately 10% open subset of an original 1.2 m…
Compress Any Segment Anything Model (SAM)
Juntong Fan, Zhiwei Hao, Jianqiang Shen +4
Due to the excellent performance in yielding high-quality, zero-shot segmentation, Segment Anything Model (SAM) and its variants have been widely applied in diverse scenarios such…
Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset
Zhuowei Chen, Bingchuan Li, Tianxiang Ma +8
Subject-to-video generation has witnessed substantial progress in recent years. However, existing models still face significant challenges in faithfully following textual instructi…