activity
20242026
most citedSridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model

1 citations · 1 across the 12 of their papers we have counts for

collaborators

18 papers

cs.HC2026

World Craft: Agentic Framework to Create Visualizable Worlds via Text

Jianwen Sun, Yukang Feng, Kaining Ying +8

Large Language Models (LLMs) motivate generative agent simulation (e.g., AI Town) to create a ``dynamic world'', holding immense value across entertainment and research. However, f…

cs.SE2025

ProSoftArena: Benchmarking Hierarchical Capabilities of Multimodal Agents in Professional Software Environments

Jiaxin Ai, Yukang Feng, Fanrui Zhang +8

Multimodal agents are making rapid progress on general computer-use tasks, yet existing benchmarks remain largely confined to browsers and basic desktop applications, falling short…

cs.CV2025

Yume-1.5: A Text-Controlled Interactive World Generation Model

Xiaofeng Mao, Zhen Li, Chuanhao Li +6

Recent approaches have demonstrated the promise of using diffusion models to generate interactive and explorable worlds. However, most of these methods face critical challenges suc…

cs.CV2025

Composition-Incremental Learning for Compositional Generalization

Zhen Li, Yuwei Wu, Chenchen Jing +3

Compositional generalization has achieved substantial progress in computer vision on pre-collected training data. Nonetheless, real-world data continually emerges, with possible co…

cs.CV2025

From Pixels to Paths: A Multi-Agent Framework for Editable Scientific Illustration

Jianwen Sun, Fanrui Zhang, Yukang Feng +6

Scientific illustrations demand both high information density and post-editability. However, current generative models have two major limitations: Frist, image generation models ou…

cs.AI2025

Multi-Step Reasoning for Embodied Question Answering via Tool Augmentation

Mingliang Zhai, Hansheng Liang, Xiaomeng Fan +6

Embodied Question Answering (EQA) requires agents to explore 3D environments to obtain observations and answer questions related to the scene. Existing methods leverage VLMs to dir…