4 papers
One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception
Keqin Zeng, Shuting Su, Shihao Lin +2
Reliable spatial decision automation, such as autonomous driving and maritime surveillance, critically depends on robust visual perception. However, real-world spatiotemporal data…
Rethinking Air-Ground Collaboration: A Progressive Cross-Task Benchmark and Socialized Learning Framework
Zhoupeng Guo, Yunqi Zhu, Zhihe Fan +6
Air-ground collaborative perception is crucial for robust visual understanding in real-world dynamic environments. However, existing studies typically formulate collaboration as si…
RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot
Liang Heng, Xiaoqi Li, Shangqing Mao +9
Recent advancements in imitation learning have shown promising results in robotic manipulation, driven by the availability of high-quality training data. To improve data collection…
ASGDiffusion: Parallel High-Resolution Generation with Asynchronous Structure Guidance
Yuming Li, Peidong Jia, Daiwei Hong +5
Training-free high-resolution (HR) image generation has garnered significant attention due to the high costs of training large diffusion models. Most existing methods begin by reco…