6 papers
AnyID: Ultra-Fidelity Universal Identity-Preserving Video Generation from Any Visual References
Jiahao Wang, Hualian Sheng, Sijia Cai +5
Identity-preserving video generation offers powerful tools for creative expression, allowing users to customize videos featuring their beloved characters. However, prevailing metho…
DepthArb: Training-Free Depth-Arbitrated Generation for Occlusion-Robust Image Synthesis
Hongjin Niu, Jiahao Wang, Xirui Hu +4
Text-to-image models often struggle to synthesize correct occlusion relationships among multiple objects, especially in densely overlapping regions. Many training-free layout-guide…
Memory-guided Prototypical Co-occurrence Learning for Mixed Emotion Recognition
Ming Li, Yong-Jin Liu, Fang Liu +6
Emotion recognition from multi-modal physiological and behavioral signals plays a pivotal role in affective computing, yet most existing models remain constrained to the prediction…
MVGSR: Multi-View Consistent 3D Gaussian Super-Resolution via Epipolar Guidance
Kaizhe Zhang, Shinan Chen, Qian Zhao +3
Scenes reconstructed by 3D Gaussian Splatting (3DGS) trained on low-resolution (LR) images are unsuitable for high-resolution (HR) rendering. Consequently, a 3DGS super-resolution…
HGS: Hybrid Gaussian Splatting with Static-Dynamic Decomposition for Compact Dynamic View Synthesis
Kaizhe Zhang, Yijie Zhou, Weizhan Zhang +5
Dynamic novel view synthesis (NVS) is essential for creating immersive experiences. Existing approaches have advanced dynamic NVS by introducing 3D Gaussian Splatting (3DGS) with i…
InfoCLIP: Bridging Vision-Language Pretraining and Open-Vocabulary Semantic Segmentation via Information-Theoretic Alignment Transfer
Muyao Yuan, Yuanhong Zhang, Weizhan Zhang +4
Recently, the strong generalization ability of CLIP has facilitated open-vocabulary semantic segmentation, which labels pixels using arbitrary text. However, existing methods that…