189 citations · 193 across the 4 of their papers we have counts for
4 papers
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Guoqing Ma, Haoyang Huang, Kun Yan +112
We present Step-Video-T2V, a state-of-the-art text-to-video pre-trained model with 30B parameters and the ability to generate videos up to 204 frames in length. A deep compression…
LayoutNUWA: Revealing the Hidden Layout Expertise of Large Language Models
Zecheng Tang, Chenfei Wu, Juntao Li +1
Graphic layout generation, a growing research field, plays a significant role in user engagement and information perception. Existing methods primarily treat layout generation as a…
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Chenfei Wu, Shengming Yin, Weizhen Qi +3
ChatGPT is attracting a cross-field interest as it provides a language interface with remarkable conversational competency and reasoning capabilities across many domains. However,…
Chinese grammatical error correction based on knowledge distillation
Peng Xia, Yuechi Zhou, Ziyan Zhang +2
In view of the poor robustness of existing Chinese grammatical error correction models on attack test sets and large model parameters, this paper uses the method of knowledge disti…