works on

From the 1 of 6 linked papers with an AI index.

activity
20242026
collaborators

6 papers

cs.CV2026

MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control

Kaiqi Liu, Yunyao Mao, Ziqi Cai +8

MAVIN is a framework for generating multi-shot audio‑visual content with fine‑grained narrative control, using boundary‑aware attention to align temporal segments and ID‑aware prop…

cs.CV2026

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation

Yiyan Xu, Qiulin Wang, Wenjie Wang +5

Multi-reference image generation aims to synthesize images from textual instructions while faithfully preserving subject identities from multiple reference images. Existing VLM-enh…

cs.LG2025

Hyper-Connections

Defa Zhu, Hongzhi Huang, Zihao Huang +5

We present hyper-connections, a simple yet effective method that can serve as an alternative to residual connections. This approach specifically addresses common drawbacks observed…

cs.CV2024

MotionRL: Align Text-to-Motion Generation to Human Preferences with Multi-Reward Reinforcement Learning

Xiaoyang Liu, Yunyao Mao, Wengang Zhou +1

We introduce MotionRL, the first approach to utilize Multi-Reward Reinforcement Learning (RL) for optimizing text-to-motion generation tasks and aligning them with human preference…

cs.CV2024

MASA: Motion-aware Masked Autoencoder with Semantic Alignment for Sign Language Recognition

Weichao Zhao, Hezhen Hu, Wengang Zhou +3

Sign language recognition (SLR) has long been plagued by insufficient model representation capabilities. Although current pre-training approaches have alleviated this dilemma to so…

cs.CV2024

Learning Generalizable Human Motion Generator with Reinforcement Learning

Yunyao Mao, Xiaoyang Liu, Wengang Zhou +2

Text-driven human motion generation, as one of the vital tasks in computer-aided content creation, has recently attracted increasing attention. While pioneering research has largel…