3 papers
cs.CV2024
Mora: Enabling Generalist Video Generation via A Multi-Agent Framework
Zhengqing Yuan, Yixin Liu, Yihan Cao +10
Text-to-video generation has made significant strides, but replicating the capabilities of advanced systems like OpenAI Sora remains challenging due to their closed-source nature.…
cs.CL2024
Instruction Mining: Instruction Data Selection for Tuning Large Language Models
Yihan Cao, Yanbin Kang, Chi Wang +1
Large language models (LLMs) are initially pretrained for broad capabilities and then finetuned with instruction-following datasets to improve their performance in interacting with…
cs.CV2024
ViT-1.58b: Mobile Vision Transformers in the 1-bit Era
Zhengqing Yuan, Rong Zhou, Hongyi Wang +3
Vision Transformers (ViTs) have achieved remarkable performance in various image classification tasks by leveraging the attention mechanism to process image patches as tokens. Howe…