14 papers
To See a World in a Living Context: Unified Indoor-Outdoor Urban World Generation
Xiaobin Huang, Zilong Huang, Yang Luo +3
Text-driven 3D generation has advanced rapidly in creating large-scale outdoor environments and detailed indoor scenes, but these domains are usually synthesized independently, lac…
WorldClaw: Agentic 3D Open-World Generation at Scale
Chunchao Guo, Jinpeng Li, Yang Li +1
Generating large-scale, freely explorable 3D worlds from open-ended text remains challenging because a system must jointly maintain global spatial coherence, rich local content, an…
GenClaw: Code-Driven Agentic Image Generation
Junyan Ye, Jun He, Zilong Huang +4
Image generation models have evolved from text-conditioned pixel synthesis toward multimodal agents endowed with visual comprehension and tool invocation capabilities. Yet, existin…
ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models
Qirui Shen, Wenda Wang, Jiachen Lu +5
Architectural spatial intelligence, the ability to recognize and infer architectural space, is fundamental to tasks such as robot navigation, embodied interaction, and 3D scene und…
SatSAM2: Motion-Constrained Video Object Tracking in Satellite Imagery using Promptable SAM2 and Kalman Priors
Ruijie Fan, Junyan Ye, Huan Chen +3
Existing satellite video tracking methods often struggle with generalization, requiring scenario-specific training to achieve satisfactory performance, and are prone to track loss…
Mind-Brush: Integrating Agentic Cognitive Search and Reasoning into Image Generation
Jun He, Junyan Ye, Zilong Huang +6
While text-to-image generation has achieved unprecedented fidelity, the vast majority of existing models function fundamentally as static text-to-pixel decoders. Consequently, they…