7 papers · 1 filter
GEAR: GEometry-motion Alternating Refinement for Articulated Object Modeling with Gaussian Splatting
Jialin Li, Bin Fu, Ruiping Wang +1
High-fidelity interactive digital assets are essential for embodied intelligence and robotic interaction, yet articulated objects remain challenging to reconstruct due to their com…
From Semantics to Pixels: Coarse-to-Fine Masked Autoencoders for Hierarchical Visual Understanding
Wenzhao Xiang, Yue Wu, Hongyang Yu +3
Self-supervised visual pre-training methods face an inherent tension: contrastive learning (CL) captures global semantics but loses fine-grained detail, while masked image modeling…
VisKnow: Constructing Visual Knowledge Base for Object Understanding
Ziwei Yao, Qiyang Wan, Ruiping Wang +1
Understanding objects is fundamental to computer vision. Beyond object recognition that provides only a category label as typical output, in-depth object understanding represents a…
A Survey on Interpretability in Visual Recognition
Qiyang Wan, Chengzhi Gao, Ruiping Wang +1
Visual recognition models have achieved unprecedented success in various tasks. While researchers aim to understand the underlying mechanisms of these models, the growing demand fo…
MoTE: Mixture of Ternary Experts for Memory-efficient Large Multimodal Models
Hongyu Wang, Jiayu Xu, Ruiping Wang +5
Large multimodal Mixture-of-Experts (MoEs) effectively scale the model size to boost performance while maintaining fixed active parameters. However, previous works primarily utiliz…
Blocks as Probes: Dissecting Categorization Ability of Large Multimodal Models
Bin Fu, Qiyang Wan, Jialin Li +2
Categorization, a core cognitive ability in humans that organizes objects based on common features, is essential to cognitive science as well as computer vision. To evaluate the ca…