activity
20202026
most citedORDNet: Capturing Omni-Range Dependencies for Scene Parsing

22 citations · 62 across the 21 of their papers we have counts for

collaborators
Showing cs.CVShow all

21 papers · 1 filter

cs.CV2026

Acoustically Grounded Cost Learning for Open-Vocabulary Audio-Visual Semantic Segmentation

Tianrui Hui, Shaofei Huang, Qisong Han +6

Open-Vocabulary Audio-Visual Semantic Segmentation (OV-AVSS) aims to perform pixel-level segmentation of sound-emitting objects from an open set of categories. The previous method…

cs.CV2026

Technical Report for the ICRA 2026 GOOSE 2D Fine-Grained Semantic Segmentation Challenge: Pretraining-Diverse Ensemble of Foundation Vision Encoders for Robust Outdoor Scene Understanding

Boyan Wang, Yongxi Huang, Wenjing Li +4

This report presents our solution for the ICRA 2026 GOOSE 2D Fine-Grained Semantic Segmentation Challenge, which requires parsing unstructured outdoor scenes from four camera platf…

cs.CV2025

RATopo: Improving Lane Topology Reasoning via Redundancy Assignment

Han Li, Shaofei Huang, Longfei Xu +3

Lane topology reasoning plays a critical role in autonomous driving by modeling the connections among lanes and the topological relationships between lanes and traffic elements. Mo…

cs.CV2025

DOMR: Establishing Cross-View Segmentation via Dense Object Matching

Jitong Liao, Yulu Gao, Shaofei Huang +4

Cross-view object correspondence involves matching objects between egocentric (first-person) and exocentric (third-person) views. It is a critical yet challenging task for visual u…

cs.CV2025

Revisiting Audio-Visual Segmentation with Vision-Centric Transformer

Shaofei Huang, Rui Ling, Tianrui Hui +6

Audio-Visual Segmentation (AVS) aims to segment sound-producing objects in video frames based on the associated audio signal. Prevailing AVS methods typically adopt an audio-centri…

cs.CV2025

LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

Hongyu Li, Jinyu Chen, Ziyu Wei +5

Recent advancements in multimodal large language models (MLLMs) have shown promising results, yet existing approaches struggle to effectively handle both temporal and spatial local…