7 papers · 1 filter
ActionSplice: In-Flight Action Editing for Interactive World Models
Pardis Taghavi, Tingyu Guo, Jonas Lossner +2
Chunk-autoregressive video world models typically condition each generated chunk on one action. An action received during sampling must therefore wait for the next chunk, condition…
Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models
Pardis Taghavi, Reza Langari, Gaurav Pandey
Training-free block-sparse attention can accelerate video transformers, but row-wise attention concentration does not by itself specify an executable sparse operator. Queries shari…
Training a Student Expert via Semi-Supervised Foundation Model Distillation
Pardis Taghavi, Tian Liu, Renjie Li +2
Foundation models deliver strong perception but are often too computationally heavy to deploy, and adapting them typically requires costly annotations. We introduce a semi-supervis…
The Pulse of Motion: Measuring Physical Frame Rate from Visual Dynamics
Xiangbo Gao, Mingyang Wu, Siyuan Yang +4
While recent generative video models have achieved remarkable visual realism and are being explored as world models, true physical simulation requires mastering both space and time…
MODEST: Multi-Optics Depth-of-Field Stereo Dataset
Nisarg K. Trivedi, Vinayaka A. Belludi, Li-Yun Wang
Training and evaluation of state-of-the-art computer vision algorithms for reliable shallow depth of field (DoF) rendering and defocus deblurring remain constrained by a persistent…
CAST: Contrastive Adaptation and Distillation for Semi-Supervised Instance Segmentation
Pardis Taghavi, Tian Liu, Renjie Li +2
Instance segmentation demands costly per-pixel annotations and computationally expensive models. We introduce CAST, a semi-supervised knowledge distillation (SSKD) framework that c…