activity
20242026
collaborators

7 papers

cs.CV2026

A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation

Yukang Feng, Jianwen Sun, Chuanhao Li +8

Recent advancements in Large Multimodal Models (LMMs) have significantly improved multimodal understanding and generation. However, these models still struggle to generate tightly…

cs.AI2026

Closing the Expression Gap in LLM Instructions via Socratic Questioning

Jianwen Sun, Yukang Feng, Yifan Chang +6

A fundamental bottleneck in human-AI collaboration is the ``intention expression gap," the difficulty for humans to effectively convey complex, high-dimensional thoughts to AI. Thi…

cs.HC2026

World Craft: Agentic Framework to Create Visualizable Worlds via Text

Jianwen Sun, Yukang Feng, Kaining Ying +8

Large Language Models (LLMs) motivate generative agent simulation (e.g., AI Town) to create a ``dynamic world'', holding immense value across entertainment and research. However, f…

cs.CV2025

From Pixels to Paths: A Multi-Agent Framework for Editable Scientific Illustration

Jianwen Sun, Fanrui Zhang, Yukang Feng +6

Scientific illustrations demand both high information density and post-editability. However, current generative models have two major limitations: Frist, image generation models ou…

cs.AI2025

BLM: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning

Wentao Tan, Bowen Wang, Heng Zhi +15

Multimodal large language models (MLLMs) have advanced vision-language reasoning and are increasingly deployed in embodied agents. However, significant limitations remain: MLLMs ge…

cs.CV2025

ARMOR: Empowering Multimodal Understanding Model with Interleaved Multimodal Generation Capability

Jianwen Sun, Yukang Feng, Chuanhao Li +7

Unified multimodal understanding and generation have recently received much attention in the area of vision and language. Existing UniMs are designed to simultaneously learn both m…