2 papers
cs.CV2026
MS-CustomNet: Controllable Multi-Subject Customization with Hierarchical Relational Semantics
Pengxiang Cai, Mengyang Li
Diffusion-based text-to-image generation has advanced significantly, yet customizing scenes with multiple distinct subjects while maintaining fine-grained control over their intera…
cs.CV2026
HiVid-Narrator: Hierarchical Video Narrative Generation with Scene-Primed ASR-anchored Compression
Haoxuan Li, Mengyan Li, Junjun Zheng
Generating structured narrations for real-world e-commerce videos requires models to perceive fine-grained visual details and organize them into coherent, high-level stories--capab…