1 citations · 1 across the 4 of their papers we have counts for
6 papers
World Craft: Agentic Framework to Create Visualizable Worlds via Text
Jianwen Sun, Yukang Feng, Kaining Ying +8
Large Language Models (LLMs) motivate generative agent simulation (e.g., AI Town) to create a ``dynamic world'', holding immense value across entertainment and research. However, f…
ProSoftArena: Benchmarking Hierarchical Capabilities of Multimodal Agents in Professional Software Environments
Jiaxin Ai, Yukang Feng, Fanrui Zhang +8
Multimodal agents are making rapid progress on general computer-use tasks, yet existing benchmarks remain largely confined to browsers and basic desktop applications, falling short…
From Pixels to Paths: A Multi-Agent Framework for Editable Scientific Illustration
Jianwen Sun, Fanrui Zhang, Yukang Feng +6
Scientific illustrations demand both high information density and post-editability. However, current generative models have two major limitations: Frist, image generation models ou…
Closing the Expression Gap in LLM Instructions via Socratic Questioning
Jianwen Sun, Yukang Feng, Yifan Chang +6
A fundamental bottleneck in human-AI collaboration is the ``intention expression gap," the difficulty for humans to effectively convey complex, high-dimensional thoughts to AI. Thi…
SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model
Yifan Chang, Yukang Feng, Jianwen Sun +4
Recent years have seen rapid advances in AI-driven image generation. Early diffusion models emphasized perceptual quality, while newer multimodal models like GPT-4o-image integrate…
IA-T2I: Internet-Augmented Text-to-Image Generation
Chuanhao Li, Jianwen Sun, Yukang Feng +3
Current text-to-image (T2I) generation models achieve promising results, but they fail on the scenarios where the knowledge implied in the text prompt is uncertain. For example, a…