5 papers
Delta-Adapter: Scalable Exemplar-Based Image Editing with Single-Pair Supervision
Jiacheng Chen, Songze Li, Han Fu +5
Exemplar-based image editing applies a transformation defined by a source-target image pair to a new query image. Existing methods rely on a pair-of-pairs supervision paradigm, req…
HiCrew: Hierarchical Reasoning for Long-Form Video Understanding via Question-Aware Multi-Agent Collaboration
Yuehan Zhu, Jingqi Zhao, Jiawen Zhao +2
Long-form video understanding remains fundamentally challenged by pervasive spatiotemporal redundancy and intricate narrative dependencies that span extended temporal horizons. Whi…
Depth-Guided Metric-Aware Temporal Consistency for Monocular Video Human Mesh Recovery
Jiaxin Cen, Xudong Mao, Guanghui Yue +4
Monocular video human mesh recovery faces fundamental challenges in maintaining metric consistency and temporal stability due to inherent depth ambiguities and scale uncertainties.…
DepthGait: Multi-Scale Cross-Level Feature Fusion of RGB-Derived Depth and Silhouette Sequences for Robust Gait Recognition
Xinzhu Li, Juepeng Zheng, Yikun Chen +7
Robust gait recognition requires highly discriminative representations, which are closely tied to input modalities. While binary silhouettes and skeletons have dominated recent lit…
VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering
Yiran Meng, Junhong Ye, Wei Zhou +4
Cross-video question answering presents significant challenges beyond traditional single-video understanding, particularly in establishing meaningful connections across video strea…