4 papers
Edit Where You Mean: Region-Aware Adapter Injection for Mask-Free Local Image Editing
Honghao Cai, Xiangyuan Wang, Yunhao Bai +8
Large diffusion transformers (DiTs) follow global editing instructions well but consistently leak local edits into unrelated regions, because joint-attention architectures offer no…
KoopmanFlow: Spectrally Decoupled Generative Control Policy via Koopman Structural Bias
Chengsi Yao, Ge Wang, Kai Kang +9
Generative Control Policies (GCPs) show immense promise in robotic manipulation but struggle to simultaneously model stable global motions and high-frequency local corrections. Whi…
SVLL: Staged Vision-Language Learning for Physically Grounded Embodied Task Planning
Yuyuan Yang, Junkun Hong, Hongrong Wang +11
Embodied task planning demands vision-language models to generate action sequences that are both visually grounded and causally coherent over time. However, existing training parad…
RoboMemory: A Brain-inspired Multi-memory Agentic Framework for Interactive Environmental Learning in Physical Embodied Systems
Mingcong Lei, Honghao Cai, Yuyuan Yang +16
Embodied intelligence aims to enable robots to learn, reason, and generalize robustly across complex real-world environments. However, existing approaches often struggle with parti…