5 papers
Arachne: Orchestrating Cascades for Efficient Text-to-Video Model Training
Peng Yu, Yuankai Fan, Yang Qiu +4
The rising demand for AI-generated videos is fueled by advances in large-scale Text-to-Video (T2V) models, trained on extensive datasets of video clips spanning diverse resolutions…
FrameOracle: Learning What to See and How Much to See in Videos
Chaoyu Li, Tianzhi Li, Fei Tao +6
Vision-language models (VLMs) advance video understanding but operate under tight computational budgets, making performance dependent on selecting a small, high-quality subset of f…
PRTS: A Primitive Reasoning and Tasking System via Contrastive Representations
Yang Zhang, Jiangyuan Zhao, Chenyou Fan +11
Vision-Language-Action (VLA) models advance robotic control via strong visual-linguistic priors. However, existing VLAs predominantly frame pretraining as supervised behavior cloni…
Tele-Omni: a Unified Multimodal Framework for Video Generation and Editing
Jialun Liu, Tian Li, Xiao Cao +20
Recent advances in diffusion-based video generation have substantially improved visual fidelity and temporal coherence. However, most existing approaches remain task-specific and r…
TeleWorld: Towards Dynamic Multimodal Synthesis with a 4D World Model
Yabo Chen, Yuanzhi Liang, Jiepeng Wang +24
World models aim to endow AI systems with the ability to represent, generate, and interact with dynamic environments in a coherent and temporally consistent manner. While recent vi…