7 papers
A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
Yukang Feng, Jianwen Sun, Chuanhao Li +8
Recent advancements in Large Multimodal Models (LMMs) have significantly improved multimodal understanding and generation. However, these models still struggle to generate tightly…
Closing the Expression Gap in LLM Instructions via Socratic Questioning
Jianwen Sun, Yukang Feng, Yifan Chang +6
A fundamental bottleneck in human-AI collaboration is the ``intention expression gap," the difficulty for humans to effectively convey complex, high-dimensional thoughts to AI. Thi…
World Craft: Agentic Framework to Create Visualizable Worlds via Text
Jianwen Sun, Yukang Feng, Kaining Ying +8
Large Language Models (LLMs) motivate generative agent simulation (e.g., AI Town) to create a ``dynamic world'', holding immense value across entertainment and research. However, f…
From Pixels to Paths: A Multi-Agent Framework for Editable Scientific Illustration
Jianwen Sun, Fanrui Zhang, Yukang Feng +6
Scientific illustrations demand both high information density and post-editability. However, current generative models have two major limitations: Frist, image generation models ou…
BLM: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
Wentao Tan, Bowen Wang, Heng Zhi +15
Multimodal large language models (MLLMs) have advanced vision-language reasoning and are increasingly deployed in embodied agents. However, significant limitations remain: MLLMs ge…
ARMOR: Empowering Multimodal Understanding Model with Interleaved Multimodal Generation Capability
Jianwen Sun, Yukang Feng, Chuanhao Li +7
Unified multimodal understanding and generation have recently received much attention in the area of vision and language. Existing UniMs are designed to simultaneously learn both m…