activity
20242026
most citedA unified multimodal understanding and generation model for cross-disciplinary scientific research

1 citations · 1 across the 2 of their papers we have counts for

collaborators

7 papers

cs.AI20261 cited

A unified multimodal understanding and generation model for cross-disciplinary scientific research

Xiaomeng Yang, Zhiyu Tan, Xiaohui Zhong +5

Scientific discovery increasingly relies on integrating heterogeneous, high-dimensional data across disciplines nowadays. While AI models have achieved notable success across vario…

cs.CV2025

RDPO: Real Data Preference Optimization for Physics Consistency Video Generation

Wenxu Qian, Chaoyue Wang, Hou Peng +3

Video generation techniques have achieved remarkable advancements in visual quality, yet faithfully reproducing real-world physics remains elusive. Preference-based model post-trai…

cs.CV2025

Cockatiel: Ensembling Synthetic and Human Preferenced Training for Detailed Video Caption

Luozheng Qin, Zhiyu Tan, Mengping Yang +2

Video Detailed Captioning (VDC) is a crucial task for vision-language bridging, enabling fine-grained descriptions of complex video content. In this paper, we first comprehensively…

cs.CV2025

SARA: Structural and Adversarial Representation Alignment for Training-efficient Diffusion Models

Hesen Chen, Junyan Wang, Zhiyu Tan +1

Modern diffusion models encounter a fundamental trade-off between training efficiency and generation quality. While existing representation alignment methods, such as REPA, acceler…

cs.CV2025

Raccoon: Multi-stage Diffusion Training with Coarse-to-Fine Curating Videos

Zhiyu Tan, Junyan Wang, Hao Yang +4

Text-to-video generation has demonstrated promising progress with the advent of diffusion models, yet existing approaches are limited by dataset quality and computational resources…

cs.CV2024

E2ED^2:Direct Mapping from Noise to Data for Enhanced Diffusion Models

Zhiyu Tan, WenXu Qian, Hesen Chen +3

Diffusion models have established themselves as the de facto primary paradigm in visual generative modeling, revolutionizing the field through remarkable success across various div…