Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
MiVE: Multiscale Vision-language features for reference-guided video Editing
Tong Wang, Meng Zou, Chengjing Wu +4
Reference-guided video editing takes a source video, a text instruction, and a reference image as inputs, requiring the model to faithfully apply the instructed edits while preserv…
cs.CV2024
Controllable Navigation Instruction Generation with Chain of Thought Prompting
Xianghao Kong, Jinyu Chen, Wenguan Wang +4
Instruction generation is a vital and multidisciplinary research area with broad applications. Existing instruction generation models are limited to generating instructions in a si…