3 papers
cs.CV2026
LeFlow: Generative Latent Flow Planning for World Models
Hsiang-Wei Huang, Jianxu Shangguan, Junbin Lu +1
Latent world models are inherently strong encoders that transform image pixel to latent embedding, yet existing world models still rely on online trajectory optimization for action…
cs.CV2026
Resonant Minds: Closed-Loop Social Avatars with Theory of Mind
Jianxu Shangguan, Jing Xu, Hang Ye +4
Creating lifelike digital humans with genuine social intelligence requires unifying cognitive reasoning and multimodal generation within a coherent framework. Current approaches tr…
cs.CV2026
CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models
Hsiang-Wei Huang, Junbin Lu, Kuang-Ming Chen +3
Vision-Language Models (VLMs) achieve strong performance on spatial question answering benchmarks, yet it remains unclear whether such gains reflect genuine spatial intelligence. W…