2 papers
cs.CV2026
PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation
Shuyan Ke, Yifan Mei, Changli Wu +4
Reasoning segmentation has recently expanded from ground-level scenes to remote-sensing imagery, yet UAV data poses distinct challenges, including oblique viewpoints, ultra-high re…
cs.CV2026
PerformRecast: Expression and Head Pose Disentanglement for Portrait Video Editing
Jiadong Liang, Bojun Xiong, Jie Tian +4
This paper primarily investigates the task of expression-only portrait video performance editing based on a driving video, which plays a crucial role in animation and film industri…