1 paper
Minghao Yin, Jiahao Lu, Wenbo Hu +3
Video diffusion transformers address their tokens by position on the pixel-time grid: an address in the tensor, not in the world. The address we would want, the world point a token…