activity
20242026
collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

MoWorld: A Flash World Model

Team Moxin, Deyi Ji, Tianrun Chen +37

The future of World Models depends not only on scaling model capability, but also on scaling practicality and inference efficiency. High-frame-rate inference enables responsive per…

cs.CV2026

MuSteerNet: Human Reaction Generation from Videos via Observation-Reaction Mutual Steering

Yuan Zhou, Yongzhi Li, Yanqi Dai +6

Video-driven human reaction generation aims to synthesize 3D human motions that directly react to observed video sequences, which is crucial for building human-like interactive AI…

cs.CV2025

CookingDiffusion: Cooking Procedural Image Generation with Stable Diffusion

Yuan Wang, Bin Zhu, Yanbin Hao +3

Recent advancements in text-to-image generation models have excelled in creating diverse and realistic images. This success extends to food imagery, where various conditional input…

cs.CV2024

Enhancing Zero-Shot Vision Models by Label-Free Prompt Distribution Learning and Bias Correcting

Xingyu Zhu, Beier Zhu, Yi Tan +3

Vision-language models, such as CLIP, have shown impressive generalization capacities when using appropriate text descriptions. While optimizing prompts on downstream labeled data…

cs.CV2024

Selective Volume Mixup for Video Action Recognition

Yi Tan, Zhaofan Qiu, Yanbin Hao +2

The recent advances in Convolutional Neural Networks (CNNs) and Vision Transformers have convincingly demonstrated high learning capability for video action recognition on large da…

cs.CV2024

Selective Vision-Language Subspace Projection for Few-shot CLIP

Xingyu Zhu, Beier Zhu, Yi Tan +3

Vision-language models such as CLIP are capable of mapping the different modality data into a unified feature space, enabling zero/few-shot inference by measuring the similarity of…