21 papers
WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN
Yuehao Huang, Yunzi Wu, Xiaotao Zhang +7
Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) into vision-language-action (VLA) policies that map egocentric observati…
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models
Paribesh Regmi, Qingshuang Chen, Chi Zhang +3
Vision-language models excel at image and video understanding but suffer from high inference latency due to the need to process thousands of tokens per image, limiting their deploy…
CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling
Yuyang Huang, Yabo Chen, Wenrui Dai +6
CineWeaver introduces a training-free method that modifies pretrained video diffusion models to generate long, multi-shot cinematic videos with fine-grained reference control and c…
ShotPlan: Cinematic Video Generation with Learnable Planning Token
Su Guo, Guangce Liu, Haosen Yang +7
Current video generation models achieve impressive results in single-shot generation, yet remain limited in cinematic video generation, where coherent narratives and effective mult…
Generative Transmission: Rethinking Computation, Bandwidth, and Memory in Communication
Xiangyu Chen, Jixiang Luo, Yuankai Fan +3
Under the AI Flow framework, communication is shifting from transmitting fidelity-oriented information flows toward delivering task-oriented and perception-oriented token flows acr…
SIFT: Self-Imagination Fine-Tuning for Physically Plausible Motion in Video Diffusion Models
Ruoyu Wang, Jialun Liu, Huayang Huang +5
The paper introduces Self-Imagination Fine-Tuning (SIFT), a method that trains video diffusion models on their own generated videos to improve physical realism and disentangle moti…