2 papers
cs.CV2026
FantasyVLN: Unified Multimodal Chain-of-Thought Reasoning for Vision-Language Navigation
Jing Zuo, Lingzhou Mu, Fan Jiang +3
Achieving human-level performance in Vision-and-Language Navigation (VLN) requires an embodied agent to jointly understand multimodal instructions and visual-spatial context while…
cs.CV2025
SkyReels-V2: Infinite-length Film Generative Model
Guibin Chen, Dixuan Lin, Jiangping Yang +22
Recent advances in video generation have been driven by diffusion models and autoregressive frameworks, yet critical challenges persist in harmonizing prompt adherence, visual qual…