collaborators

10 papers

cs.CV2026

SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation

Longtao Jiang, Jianmin Bao, Zhendong Wang +4

Normalizing flows (NFs) provide exact likelihoods and deterministic invertible sampling, but have historically lagged behind diffusion models for large-scale image generation. We i…

cs.CV2026

LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning

Haihong Hao, Lei Chen, Mingfei Han +5

Existing vision-and-language navigation (VLN) models primarily reason over past and current visual observations, while largely ignoring the future visual dynamics induced by action…

cs.RO2026

Beyond Dense Futures: World Models as Structured Planners for Robotic Manipulation

Minghao Jin, Mozheng Liao, Mingfei Han +2

Recent world-model-based Vision-Language-Action (VLA) architectures have improved robotic manipulation through predictive visual foresight. However, dense future prediction introdu…

cs.AI2026

User-Feedback-Driven Adaptation for Vision-and-Language Navigation

Yongqiang Yu, Xuhui Li, Hazza Mahmood +6

Real-world deployment of Vision-and-Language Navigation (VLN) agents is constrained by the scarcity of reliable supervision after offline training. While recent adaptation methods…

cs.CV2025

Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive Learning

Changlin Li, Jiawei Zhang, Shuhao Liu +4

Human video generation has advanced rapidly with the development of diffusion models, but the high computational cost and substantial memory consumption associated with training th…

cs.CV2025

Which Layer Causes Distribution Deviation? Entropy-Guided Adaptive Pruning for Diffusion and Flow Models

Changlin Li, Jiawei Zhang, Zeyi Shi +3

Large-scale vision generative models, including diffusion and flow models, have demonstrated remarkable performance in visual generation tasks. However, transferring these pre-trai…