2 citations · 2 across the 4 of their papers we have counts for
4 papers · 1 filter
Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation
Jiangning Zhang, Junwei Zhu, Zhenye Gan +14
We propose a multimodal-driven framework for high-fidelity long-term digital human animation termed , which generates semantically coherent videos from a single-fram…
LLaVA-VSD: Large Language-and-Vision Assistant for Visual Spatial Description
Yizhang Jin, Jian Li, Jiangning Zhang +7
Visual Spatial Description (VSD) aims to generate texts that describe the spatial relationships between objects within images. Traditional visual spatial relationship classificatio…
PSPU: Enhanced Positive and Unlabeled Learning by Leveraging Pseudo Supervision
Chengjie Wang, Chengming Xu, Zhenye Gan +3
Positive and Unlabeled (PU) learning, a binary classification model trained with only positive and unlabeled data, generally suffers from overfitted risk estimation due to inconsis…
DMAD: Dual Memory Bank for Real-World Anomaly Detection
Jianlong Hu, Xu Chen, Zhenye Gan +7
Training a unified model is considered to be more suitable for practical industrial anomaly detection scenarios due to its generalization ability and storage efficiency. However, t…