5 papers
World Craft: Agentic Framework to Create Visualizable Worlds via Text
Jianwen Sun, Yukang Feng, Kaining Ying +8
Large Language Models (LLMs) motivate generative agent simulation (e.g., AI Town) to create a ``dynamic world'', holding immense value across entertainment and research. However, f…
From Pixels to Paths: A Multi-Agent Framework for Editable Scientific Illustration
Jianwen Sun, Fanrui Zhang, Yukang Feng +6
Scientific illustrations demand both high information density and post-editability. However, current generative models have two major limitations: Frist, image generation models ou…
BLM: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
Wentao Tan, Bowen Wang, Heng Zhi +15
Multimodal large language models (MLLMs) have advanced vision-language reasoning and are increasingly deployed in embodied agents. However, significant limitations remain: MLLMs ge…
Closing the Expression Gap in LLM Instructions via Socratic Questioning
Jianwen Sun, Yukang Feng, Yifan Chang +6
A fundamental bottleneck in human-AI collaboration is the ``intention expression gap," the difficulty for humans to effectively convey complex, high-dimensional thoughts to AI. Thi…
ARMOR: Empowering Multimodal Understanding Model with Interleaved Multimodal Generation Capability
Jianwen Sun, Yukang Feng, Chuanhao Li +7
Unified multimodal understanding and generation have recently received much attention in the area of vision and language. Existing UniMs are designed to simultaneously learn both m…