6 citations · 6 across the 10 of their papers we have counts for
4 papers · 1 filter
ST-Gen4D: Embedding 4D Spatiotemporal Cognition into World Model for 4D Generation
Haonan Wang, Hanyu Zhou, Tao Gu +1
Generative models have achieved success in producing apparently coherent 2D videos, but remain challenging in the physical world due to lack of 4D spatiotemporal scale. Typically,…
Lost in Adaptation: Layer-Selective Recovery of Temporal Reasoning in Video-Language Models
Zihang Fu, Haonan Wang, Jian Kang +2
Multimodal adaptation can erode temporal reasoning (TR) in video-language models (VLMs), leaving models able to perceive salient events yet unable to infer their temporal and causa…
Cog2Gen3D: Sculpturing 3D Semantic-Geometric Cognition for 3D Generation
Haonan Wang, Hanyu Zhou, Haoyue Liu +2
Generative models have achieved success in producing semantically plausible 2D images, but it remains challenging in 3D generation due to the absence of spatial geometry constraint…
Getting More Juice Out of Your Data: Hard Pair Refinement Enhances Visual-Language Models Without Extra Data
Haonan Wang, Minbin Huang, Runhui Huang +7
Contrastive Language-Image Pre-training (CLIP) has become the standard for cross-modal image-text representation learning. Improving CLIP typically requires additional data and ret…