2 papers
cs.CV2025
Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation
Jiangning Zhang, Junwei Zhu, Zhenye Gan +14
We propose a multimodal-driven framework for high-fidelity long-term digital human animation termed , which generates semantically coherent videos from a single-fram…
cs.CV2024
LLaVA-VSD: Large Language-and-Vision Assistant for Visual Spatial Description
Yizhang Jin, Jian Li, Jiangning Zhang +7
Visual Spatial Description (VSD) aims to generate texts that describe the spatial relationships between objects within images. Traditional visual spatial relationship classificatio…