5 papers
Video Understanding with Large Language Models: A Survey
Yolo Y. Tang, Jing Bi, Siting Xu +17
With the burgeoning growth of online video platforms and the escalating volume of video content, the demand for proficient video understanding tools has intensified markedly. Given…
Latent-Reframe: Enabling Camera Control for Video Diffusion Model without Training
Zhenghong Zhou, Jie An, Jiebo Luo
Precise camera pose control is crucial for video generation with diffusion models. Existing methods require fine-tuning with additional datasets containing paired videos and camera…
On Inductive Biases That Enable Generalization of Diffusion Transformers
Jie An, De Wang, Pengsheng Guo +2
Recent work studying the generalization of diffusion models with UNet-based denoisers reveals inductive biases that can be expressed via geometry-adaptive harmonic bases. However,…
MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models
Hang Hua, Yunlong Tang, Ziyun Zeng +5
The advent of large Vision-Language Models (VLMs) has significantly advanced multimodal understanding, enabling more sophisticated and accurate integration of visual and textual in…
GaussianStyle: Gaussian Head Avatar via StyleGAN
Pinxin Liu, Luchuan Song, Daoan Zhang +5
Existing methods like Neural Radiation Fields (NeRF) and 3D Gaussian Splatting (3DGS) have made significant strides in facial attribute control such as facial animation and compone…