5 papers · 1 filter
MoWorld: A Flash World Model
Team Moxin, Deyi Ji, Tianrun Chen +37
The future of World Models depends not only on scaling model capability, but also on scaling practicality and inference efficiency. High-frame-rate inference enables responsive per…
NOUS: Video-Driven 3D Human Reaction Generation via Observation-Reaction Mutual Steering
Yuan Zhou, Luanyuan Dai, Yongzhi Li +7
Video-driven 3D human reaction generation aims to synthesize 3D human motion in response to the action observed in a video, playing an important role in interactive multimedia syst…
CookingDiffusion: Cooking Procedural Image Generation with Stable Diffusion
Yuan Wang, Bin Zhu, Yanbin Hao +3
Recent advancements in text-to-image generation models have excelled in creating diverse and realistic images. This success extends to food imagery, where various conditional input…
Enhancing Zero-Shot Vision Models by Label-Free Prompt Distribution Learning and Bias Correcting
Xingyu Zhu, Beier Zhu, Yi Tan +3
Vision-language models, such as CLIP, have shown impressive generalization capacities when using appropriate text descriptions. While optimizing prompts on downstream labeled data…
Selective Vision-Language Subspace Projection for Few-shot CLIP
Xingyu Zhu, Beier Zhu, Yi Tan +3
Vision-language models such as CLIP are capable of mapping the different modality data into a unified feature space, enabling zero/few-shot inference by measuring the similarity of…