collaborators

13 papers

cs.CV2026

Human-Centric Intelligence in the Era of Foundation Models: A Survey

Yang Chen, Tianqi Wang, Xiaorui Jiang +13

Human-centric intelligence is evolving in the foundation-model era, with growing emphasis on scale, transferability, and general-purpose modeling. Yet it has not fully integrated w…

cs.CV2026

GrainGS: Gradient-Decoupled Gaussian Splatting for Efficient Dynamic Novel View Synthesis

Jiahao He, Yihua Shao, Zhengkai Zhao +6

Dynamic scene reconstruction with 3D Gaussian Splatting requires a balance between fine-grained motion modeling, structural stability, and compact representation. Existing per-prim…

cs.CV2026

DepthART: Scaling Foundation Monocular Depth to Tiny Models

Feng Xue, Wu Chen, Mingshuai Zhao +7

Recent geometric foundation models (e.g., Metric3D, Depth Anything and UniDepth) have substantially improved monocular depth estimation (MDE) in both cross-scene generalization and…

cs.CV2026

Dress-ED: Instruction-Guided Editing for Virtual Try-On and Try-Off

Davide Lobba, Fulvio Sanguigni, Bin Ren +3

Recent advances in Virtual Try-On (VTON) and Virtual Try-Off (VTOFF) have greatly improved photo-realistic fashion synthesis and garment reconstruction. However, existing datasets…

cs.CV2026

GenHOI: Generalized Hand-Object Pose Estimation with Occlusion Awareness

Hui Yang, Wei Sun, Jian Liu +6

Generalized 3D hand-object pose estimation from a single RGB image remains challenging due to the large variations in object appearances and interaction patterns, especially under…

cs.CV2026

OmniCoT: A Benchmark for Global and Multi-Step Panoramic Reasoning

Haocong He, Chenfei Liao, Zichen Wen +13

Multimodal Large Language Models (MLLMs) have demonstrated promising spatial reasoning capabilities, while these abilities remain underexplored in the emerging visual modality of p…