8 papers
POPS: Recovering Unlearned Multi-Modality Knowledge in MLLMs with Prompt-Optimized Parameter Shaking
Zhangheng LI, Jianing Zhu, Junyuan Hong +4
Multimodal Large Language Models (MLLMs) have demonstrated impressive performance on cross-modal tasks by jointly training on large-scale textual and visual data, where privacy-sen…
Coarse-to-Fine Hierarchical Alignment for UAV-based Human Detection using Diffusion Models
Wenda Li, Meng Wu, Liangzhao Chen +3
Training object detectors demands extensive, task-specific annotations, yet this requirement becomes impractical in UAV-based human detection due to constantly shifting target dist…
Generative Spatiotemporal Data Augmentation
Jinfan Zhou, Lixin Luo, Sungmin Eum +2
We explore spatiotemporal data augmentation using video foundation models to diversify both camera viewpoints and scene dynamics. Unlike existing approaches based on simple geometr…
SynPlay: Large-Scale Synthetic Human Data with Real-World Diversity for Aerial-View Perception
Jinsub Yim, Hyungtae Lee, Sungmin Eum +4
We introduce SynPlay, a large-scale synthetic human dataset purpose-built for advancing multi-perspective human localization, with a predominant focus on aerial-view perception. Sy…
MoRe: Monocular Geometry Refinement via Graph Optimization for Cross-View Consistency
Dongki Jung, Jaehoon Choi, Yonghan Lee +3
Monocular 3D foundation models offer an extensible solution for perception tasks, making them attractive for broader 3D vision applications. In this paper, we propose MoRe, a train…
AutoComPose: Automatic Generation of Pose Transition Descriptions for Composed Pose Retrieval Using Multimodal LLMs
Yi-Ting Shen, Sungmin Eum, Doheon Lee +4
Composed pose retrieval (CPR) enables users to search for human poses by specifying a reference pose and a transition description, but progress in this field is hindered by the sca…