From the 1 of 9 linked papers with an AI index.
9 papers
From Dense Prediction to Visual Editing: Structured Supervision for Unified Image and Video Creation
Zhefan Rao, Bin Zou, Haoxuan Che +5
Unified image and video creation requires a model to follow diverse instructions while preserving identity, geometry, and temporal structure from visual context. However, semantic-…
OPTD: On-Policy Transition Distillation with Consistency-Guided Adaptive Compression for Few-Step Diffusion Language Models
Xiaocheng Lu, Hualei Zhang, Shuhan Guo +8
Diffusion language models (dLLMs) can predict many tokens in parallel, but accurate generation still requires many iterative denoising steps. Few-step distillation accelerates deco…
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget
Guoxuan Chen, Chufeng Xiao, Haoran Yang +30
Boogu-Image-0.1 is an open-source multimodal model family that supports high-quality text-to-image generation, fast inference, instruction-based image editing, and bilingual (Chine…
Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond
Meng Chu, Xuan Billy Zhang, Kevin Qinghong Lin +47
As AI systems move from generating text to accomplishing goals through sustained interaction, the ability to model environment dynamics becomes a central bottleneck. Agents that ma…
UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion
Zhaoqing Li, Haoning Xu, Jingran Su +9
We present UNISON, a latent diffusion framework that unifies speech generation, sound generation, and audio editing within a single model. A single model handles text-to-audio, tex…
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm
Yaofang Liu, Kangning Cui, Meng Chu +7
Humans often specify and create through visual artifacts: typography sheets, sketches, reference images, and annotated scenes. Yet modern visual generators still ask users to seria…