collaborators

8 papers

cs.CV2025

Binary-Gaussian: Compact and Progressive Representation for 3D Gaussian Segmentation

An Yang, Chenyu Liu, Jun Du +6

3D Gaussian Splatting (3D-GS) has emerged as an efficient 3D representation and a promising foundation for semantic tasks like segmentation. However, existing 3D-GS-based segmentat…

cs.SD2025

HDA-SELD: Hierarchical Cross-Modal Distillation with Multi-Level Data Augmentation for Low-Resource Audio-Visual Sound Event Localization and Detection

Qing Wang, Ya Jiang, Hang Chen +3

This work presents HDA-SELD, a unified framework that combines hierarchical cross-modal distillation (HCMD) and multi-level data augmentation to address low-resource audio-visual (…

cs.SD2025

Improving Anomalous Sound Detection with Attribute-aware Representation from Domain-adaptive Pre-training

Xin Fang, Guirui Zhong, Qing Wang +7

Anomalous Sound Detection (ASD) is often formulated as a machine attribute classification task, a strategy necessitated by the common scenario where only normal data is available f…

cs.GR2025

READ: Real-time and Efficient Asynchronous Diffusion for Audio-driven Talking Head Generation

Haotian Wang, Yuzhe Weng, Jun Du +7

The introduction of diffusion models has brought significant advances to the field of audio-driven talking head generation. However, the extremely slow inference speed severely lim…

cs.CL2025

Enhancing the Geometric Problem-Solving Ability of Multimodal LLMs via Symbolic-Neural Integration

Yicheng Pan, Zhenrong Zhang, Pengfei Hu +6

Recent advances in Multimodal Large Language Models (MLLMs) have achieved remarkable progress in general domains and demonstrated promise in multimodal mathematical reasoning. Howe…

cs.MM2025

MMC: Iterative Refinement of VLM Reasoning via MCTS-based Multimodal Critique

Shuhang Liu, Zhenrong Zhang, Pengfei Hu +7

Visual language models (VLMs) have demonstrated strong performance across diverse multimodal reasoning tasks but still face challenges such as hallucinations, resulting in incorrec…