11 papers
MetaPoint: Unlocking Precise Spatial Control in Agentic Visual Generation
Dewei Zhou, Xinyu Huang, Xun Wang +6
Generative visual models fundamentally struggle with precise spatial control. This arises from a core disconnect: models can process textual descriptions of space but cannot direct…
Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation
Yabo Zhang, Kunchang Li, Dewei Zhou +2
While recent advancements in multimodal language models have enabled image generation from expressive multi-image instructions, existing methods struggle to maintain performance un…
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts
You Li, Dewei Zhou, Fan Ma +3
Recent Video-to-Audio (V2A) methods have achieved remarkable progress, enabling the synthesis of realistic, high-quality audio. However, they struggle with fine-grained temporal co…
PhyEdit: Towards Real-World Object Manipulation via Physically-Grounded Image Editing
Ruihang Xu, Dewei Zhou, Xiaolong Shen +2
Achieving physically accurate object manipulation in image editing is essential for its potential applications in interactive world models. However, existing visual generative mode…
RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details
Dewei Zhou, You Li, Zongxin Yang +1
We introduce region-specific image refinement as a dedicated problem setting: given an input image and a user-specified region (e.g., a scribble mask or a bounding box), the goal i…
ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance Generation
Ruihang Xu, Dewei Zhou, Fan Ma +1
Multi-instance image generation (MIG) remains a significant challenge for modern diffusion models due to key limitations in achieving precise control over object layout and preserv…