activity
20242026
collaborators

5 papers

cs.CV2026

Mural: Transferring LLM knowledge to image generation via Mixture-of-Transformers

Achin Jain, Jie An, Siddharth Chaudhary +1

Leveraging capabilities of large language models (LLMs) in text-to-image (T2I) synthesis is an important research direction. In this work we investigate whether the knowledge of a…

cs.CV2025

Video Understanding with Large Language Models: A Survey

Yolo Y. Tang, Jing Bi, Siting Xu +17

With the burgeoning growth of online video platforms and the escalating volume of video content, the demand for proficient video understanding tools has intensified markedly. Given…

cs.CV2025

Ouroboros-Diffusion: Exploring Consistent Content Generation in Tuning-free Long Video Diffusion

Jingyuan Chen, Fuchen Long, Jie An +4

The first-in-first-out (FIFO) video diffusion, built on a pre-trained text-to-video model, has recently emerged as an effective approach for tuning-free long video generation. This…

cs.CV2024

Latent-Reframe: Enabling Camera Control for Video Diffusion Model without Training

Zhenghong Zhou, Jie An, Jiebo Luo

Precise camera pose control is crucial for video generation with diffusion models. Existing methods require fine-tuning with additional datasets containing paired videos and camera…

cs.CV2024

On Inductive Biases That Enable Generalization of Diffusion Transformers

Jie An, De Wang, Pengsheng Guo +2

Recent work studying the generalization of diffusion models with UNet-based denoisers reveals inductive biases that can be expressed via geometry-adaptive harmonic bases. However,…