activity
20242026
most citedConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning

1 citations · 3 across the 54 of their papers we have counts for

collaborators
Showing cs.CVShow all

55 papers · 1 filter

cs.CV2026

HandsOnWorld: Unconstrained Egocentric Video Generation with Camera-Disentangled Hand Control

Yushuo Chen, Xiaoyu Shi, Xiaoshi Wu +3

We present HandsOnWorld, a framework for hand-controlled egocentric video generation that learns directly from unconstrained monocular video. Prior generators depend on 3D hand ann…

cs.CV2026

ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing

Xu Guo, Zhengxuan Wei, Xinghui Li +11

Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs.…

cs.CV2026

Temporal Concentration from Rollout Errors: Implicit Preference Optimization for Text-to-Video Diffusion

Henglin Liu, Fangyuan Kong, Jing Wang +7

Recent advances in preference alignment for diffusion-based video generation, particularly via Direct Preference Optimization (DPO), have significantly improved visual quality. How…

cs.CV2026

Vera: Identity-Faithful Human Subject-to-Video Generation

Yulong Xu, Xinyue Liu, Shujuan Li +6

Subject-to-video (S2V) generation has made substantial progress in preserving reference subjects across diverse categories, yet generic subject consistency remains insufficient for…

cs.CV2026

MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control

Kaiqi Liu, Yunyao Mao, Ziqi Cai +8

While recent generative models produce high-fidelity videos, they struggle with the complex narrative control required for coherent multi-shot audio-visual generation. Existing met…

cs.CV2026

MemLearner: Learning to Query Context memory for Video World Models

Jiwen Yu, Jianxiong Gao, Jianhong Bai +7

Video World Models are interactive video generation models that predict future world states based on user actions and history video frames. A critical challenge in video world mode…