8 papers · 1 filter
APT: Atomic Physical Transitions for Causal Video-Language Understanding
Shang Wu, Haoran Lu, Songling Liu +6
Physical events are not understood by their names alone, but by the causal state changes that compose them. A clip-level label such as "bounce" can be correct while hiding the proc…
CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies
Fan Du, Feng Yan, Jianxiong Wu +8
Flow-based vision-language-action (VLA) policies offer strong expressivity for action generation, but suffer from a fundamental inefficiency: multi-step inference is required to re…
PhyPrompt: RL-based Prompt Refinement for Physically Plausible Text-to-Video Generation
Shang Wu, Chenwei Xu, Zhuofan Xia +6
State-of-the-art text-to-video (T2V) generators frequently violate physical laws despite high visual quality. We show this stems from insufficient physical constraints in prompts r…
Phys4D: Fine-Grained Physics-Consistent 4D Modeling from Video Diffusion
Haoran Lu, Shang Wu, Songling Liu +9
Recent video diffusion models have achieved impressive capabilities as large-scale generative world models. However, these models often struggle with fine-grained physical consiste…
Towards Sparse Video Understanding and Reasoning
Chenwei Xu, Zhen Ye, Shang Wu +8
We present \revise (\underline{Re}asoning with \underline{Vi}deo \underline{S}parsity), a multi-round agent for video question answering (VQA). Instead of uniformly sampling frames…
DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
Divyansh Srivastava, Akshay Mehra, Pranav Maneriker +5
Decoder-only autoregressive image generation typically relies on fixed-length tokenization schemes whose token counts grow quadratically with resolution, substantially increasing t…