5 papers
MoWorld: A Flash World Model
Team Moxin, Deyi Ji, Tianrun Chen +37
The future of World Models depends not only on scaling model capability, but also on scaling practicality and inference efficiency. High-frame-rate inference enables responsive per…
MuSteerNet: Human Reaction Generation from Videos via Observation-Reaction Mutual Steering
Yuan Zhou, Yongzhi Li, Yanqi Dai +6
Video-driven human reaction generation aims to synthesize 3D human motions that directly react to observed video sequences, which is crucial for building human-like interactive AI…
CookingDiffusion: Cooking Procedural Image Generation with Stable Diffusion
Yuan Wang, Bin Zhu, Yanbin Hao +3
Recent advancements in text-to-image generation models have excelled in creating diverse and realistic images. This success extends to food imagery, where various conditional input…
Enhancing Zero-Shot Vision Models by Label-Free Prompt Distribution Learning and Bias Correcting
Xingyu Zhu, Beier Zhu, Yi Tan +3
Vision-language models, such as CLIP, have shown impressive generalization capacities when using appropriate text descriptions. While optimizing prompts on downstream labeled data…
Selective Volume Mixup for Video Action Recognition
Yi Tan, Zhaofan Qiu, Yanbin Hao +2
The recent advances in Convolutional Neural Networks (CNNs) and Vision Transformers have convincingly demonstrated high learning capability for video action recognition on large da…