3 papers
cs.CV2025
IC-Effect: Precise and Efficient Video Effects Editing via In-Context Learning
Yuanhang Li, Yiren Song, Junzhe Bai +4
We propose \textbf{IC-Effect}, an instruction-guided, DiT-based framework for few-shot video VFX editing that synthesizes complex effects (\eg flames, particles and cartoon charact…
cs.CV2025
Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
Yuxin Song, Wenkai Dong, Shizun Wang +8
Unified Multimodal Models (UMMs) have demonstrated remarkable performance in text-to-image generation (T2I) and editing (TI2I), whether instantiated as assembled unified frameworks…
cs.RO2025
AIRoA MoMa Dataset: A Large-Scale Hierarchical Dataset for Mobile Manipulation
Ryosuke Takanami, Petr Khrapchenkov, Shu Morikuni +32
As robots transition from controlled settings to unstructured human environments, building generalist agents that can reliably follow natural language instructions remains a centra…