activity
20232026
collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

SCAPO: Self-Supervised Category-Level Articulated Pose Estimation from a Single 3D Observation

Can Zhang, Gim Hee Lee

Existing methods for category-level object articulation from a single 3D observation often rely on dense supervision, multi-frame inputs, or CAD templates, and still struggle to di…

cs.CV2025

CLAIR: CLIP-Aided Weakly Supervised Zero-Shot Cross-Domain Image Retrieval

Chor Boon Tan, Conghui Hu, Gim Hee Lee

The recent growth of large foundation models that can easily generate pseudo-labels for huge quantity of unlabeled data makes unsupervised Zero-Shot Cross-Domain Image Retrieval (U…

cs.CV2025

IAAO: Interactive Affordance Learning for Articulated Objects in 3D Environments

Can Zhang, Gim Hee Lee

This work presents IAAO, a novel framework that builds an explicit 3D model for intelligent agents to gain understanding of articulated objects in their environment through interac…

cs.CV2025

econSG: Efficient and Multi-view Consistent Open-Vocabulary 3D Semantic Gaussians

Can Zhang, Gim Hee Lee

The primary focus of most recent works on open-vocabulary neural fields is extracting precise semantic features from the VLMs and then consolidating them efficiently into a multi-v…

cs.CV2024

MVSDet: Multi-View Indoor 3D Object Detection via Efficient Plane Sweeps

Yating Xu, Chen Li, Gim Hee Lee

The key challenge of multi-view indoor 3D object detection is to infer accurate geometry information from images for precise 3D detection. Previous method relies on NeRF for geomet…

cs.CV2023

Rethink Cross-Modal Fusion in Weakly-Supervised Audio-Visual Video Parsing

Yating Xu, Conghui Hu, Gim Hee Lee

Existing works on weakly-supervised audio-visual video parsing adopt hybrid attention network (HAN) as the multi-modal embedding to capture the cross-modal context. It embeds the a…