3 papers
cs.AI2026
UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World Model
Changxin Huang, Lv Tang, Zhaohuan Zhan +5
Vision-and-Language Navigation (VLN) requires agents to autonomously navigate complex environments via visual images and natural language instructions--remains highly challenging.…
cs.LG2025
Personalized Subgraph Federated Learning with Differentiable Auxiliary Projections
Wei Zhuo, Zhaohuan Zhan, Han Yu
Federated Learning (FL) on graph-structured data typically faces non-IID challenges, particularly in scenarios where each client holds a distinct subgraph sampled from a global gra…
cs.CV2025
HouseTune: Two-Stage Floorplan Generation with LLM Assistance
Ziyang Zong, Guanying Chen, Zhaohuan Zhan +2
This paper proposes a two-stage text-to-floorplan generation framework that combines the reasoning capability of Large Language Models (LLMs) with the generative power of diffusion…