From the 2 of 10 linked papers with an AI index.
10 papers
Training Skills Like Parameters via Self-Supervised Semantic Diffusion
Mo Li, Zixin Yin, Ting Cao +1
The paper introduces a self‑supervised framework that lets a language model acquire and store textual skills in an external library using diffusion‑style reconstruction loss, witho…
Motion4Motion: Motion Transfer Across Subjects at Inference
Ling-Hao Chen, Zixin Yin, Duomin Wang +2
The paper introduces Motion4Motion, a training‑free framework that transfers motion between videos by modeling motion flow instead of relying on predefined skeletons, enabling tran…
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Haiwen Diao, Penghao Wu, Hanming Deng +55
Recent large vision-language models (VLMs) remain fundamentally constrained by a persistent dichotomy: understanding and generation are treated as distinct problems, leading to fra…
LazyDrag: Enabling Stable Drag-Based Editing on Multi-Modal Diffusion Transformers via Explicit Correspondence
Zixin Yin, Xili Dai, Duomin Wang +4
The reliance on implicit point matching via attention has become a core bottleneck in drag-based editing, resulting in a fundamental compromise on weakened inversion strength and c…
Training-Free Text-Guided Color Editing with Multi-Modal Diffusion Transformer
Zixin Yin, Xili Dai, Ling-Hao Chen +7
Text-guided color editing in images and videos is a fundamental yet unsolved problem, requiring fine-grained manipulation of color attributes, including albedo, light source color,…
ReasonEdit: Towards Reasoning-Enhanced Image Editing Models
Fukun Yin, Shiyu Liu, Yucheng Han +12
Recent advances in image editing models have shown remarkable progress. A common architectural design couples a multimodal large language model (MLLM) encoder with a diffusion deco…