21 papers
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents
Yunlong Lin, Zixu Lin, Zhaohu Xing +23
Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent generative models can synthesize high-quality images, videos, aud…
Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models
Shangwen Zhu, Qianyu Peng, Zhao Pu +12
Modern interactive video world models have achieved impressive visual fidelity, yet lack fine-grained multi-entity control and cross-entity, cross-world generalization. We trace th…
SegDINO: Introducing Multi-Scale Structure into DINO for Efficient Medical Image Segmentation
Sicheng Yang, Hongqiu Wang, Zhaohu Xing +5
Self-supervised DINO models provide strong transferable visual representations, yet applying them directly to image segmentation remains challenging. Existing approaches commonly r…
SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models
Zizhao Tong, Yeying Jin, Hongfeng Lai +11
Interactive world models for first-person shooter (FPS) games must resolve high-frequency overlapping control signals at every frame without disrupting unaffected regions. Existing…
EchoPilot: Training-Free Ultrasound Video Segmentation via Scale-Space Semantic Prompting and Reliability-Gated Memory
Ruiqiang Xiao, Zhaohu Xing, Yijun Yang +4
Ultrasound video segmentation is clinically valuable yet difficult due to speckle noise, weak boundaries, and rapid anatomical deformation. Recent promptable foundation models enab…
GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation
Sixiang Chen, Zhaohu Xing, Tian Ye +7
Open-ended image generation is no longer a simple prompt-to-image problem. High-quality generation often requires an agent to combine a model's internal generative ability with ext…