4 papers
DrawAI: Agentic Benchmark and Workflow for Making Raster Images Editable
Pu Cao, Qingye Kong, Xuedan Yin +5
Recent image-generation models and multimodal agents can produce high-quality visuals for increasingly complex visual communication tasks. Yet their raster outputs remain difficult…
Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation
Guo Ye, Zexi Zhang, Xu Zhao +4
Vision-Language-Action (VLA) models have shown remarkable generalization by mapping web-scale knowledge to robotic control, yet they remain blind to physical contact. Consequently,…
DreamOmni3: Scribble-based Editing and Generation
Bin Xia, Bohao Peng, Jiyang Liu +8
Recently unified generation and editing models have achieved remarkable success with their impressive performance. These models rely mainly on text prompts for instruction-based ed…
Preliminary Explorations with GPT-4o(mni) Native Image Generation
Pu Cao, Feng Zhou, Junyi Ji +8
Recently, the visual generation ability by GPT-4o(mni) has been unlocked by OpenAI. It demonstrates a very remarkable generation capability with excellent multimodal condition unde…