5 papers
GigaWorld-0: World Models as Data Engine to Empower Embodied AI
GigaWorld Team, Angen Ye, Boyuan Wang +22
World models are emerging as a foundational paradigm for scalable, data-efficient embodied AI. In this work, we present GigaWorld-0, a unified world model framework designed explic…
\textsc{GUI-Spotlight}: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding
Bin Lei, Nuo Xu, Ali Payani +4
Multimodal large language models (MLLMs) have markedly expanded the competence of graphical user-interface (GUI) systems, propelling them beyond controlled simulations into complex…
See What You Need: Query-Aware Visual Intelligence through Reasoning-Perception Loops
Zixuan Dong, Baoyun Peng, Yufei Wang +4
Human video comprehension demonstrates dynamic coordination between reasoning and visual attention, adaptively focusing on query-relevant details. However, current long-form video…
EmbodiedGen: Towards a Generative 3D World Engine for Embodied Intelligence
Xinjie Wang, Liu Liu, Yu Cao +5
Constructing a physically realistic and accurately scaled simulated 3D world is crucial for the training and evaluation of embodied intelligence tasks. The diversity, realism, low…
Hand1000: Generating Realistic Hands from Text with Only 1,000 Images
Haozhuo Zhang, Bin Zhu, Yu Cao +1
Text-to-image generation models have achieved remarkable advancements in recent years, aiming to produce realistic images from textual descriptions. However, these models often str…