papers

Publications (35)

cs.CV2023

Focus on Local Regions for Query-based Object Detection

Hongbin Xu, Yamei Xia, Shuai Zhao +1

Query-based methods have garnered significant attention in object detection since the advent of DETR, the pioneering query-based detector. However, these methods face challenges li…

cs.CV2021

Digging into Uncertainty in Self-supervised Multi-view Stereo

Hongbin Xu, Zhipeng Zhou, Yali Wang +4

Self-supervised Multi-view stereo (MVS) with a pretext task of image reconstruction has achieved significant progress recently. However, previous methods are built upon intuitions,…

cs.CV2023

A Simple Framework for 3D Occupancy Estimation in Autonomous Driving

Wanshui Gan, Ningkai Mo, Hongbin Xu +1

The task of estimating 3D occupancy from surrounding-view images is an exciting development in the field of autonomous driving, following the success of Bird's Eye View (BEV) perce…

cs.CV2026

Resolving Long-Tail Ambiguity in Unsupervised 3D Point Cloud Segmentation with Language Priors

Siqi Wei, Hongbin Xu, Feng Xiao +4

Existing approaches for unsupervised 3D point cloud segmentation predominantly rely on a purely visual similarity-based learning-by-clustering paradigm, which suffers from a fundam…

cs.CV2025

Safe-Sora: Safe Text-to-Video Generation via Graphical Watermarking

Zihan Su, Xuerui Qiu, Hongbin Xu +6

The explosive growth of generative video models has amplified the demand for reliable copyright preservation of AI-generated content. Despite its popularity in image synthesis, inv…

cs.CV2024

V4d: voxel for 4d novel view synthesis

Wanshui Gan, Hongbin Xu, Yi Huang +2

Neural radiance fields have made a remarkable breakthrough in the novel view synthesis task at the 3D static scene. However, for the 4D circumstance (e.g., dynamic scene), the perf…

cs.CR2025

OptMark: Robust Multi-bit Diffusion Watermarking via Inference Time Optimization

Jiazheng Xing, Hai Ci, Hongbin Xu +3

Watermarking diffusion-generated images is crucial for copyright protection and user tracking. However, current diffusion watermarking methods face significant limitations: zero-bi…

cs.CV2024

RobustMVS: Single Domain Generalized Deep Multi-view Stereo

Hongbin Xu, Weitao Chen, Baigui Sun +2

Despite the impressive performance of Multi-view Stereo (MVS) approaches given plenty of training samples, the performance degradation when generalizing to unseen domains has not b…

cs.CV2025

B2N3D: Progressive Learning from Binary to N-ary Relationships for 3D Object Grounding

Feng Xiao, Hongbin Xu, Hai Ci +1

Localizing 3D objects using natural language is essential for robotic scene understanding. The descriptions often involve multiple spatial relationships to distinguish similar obje…

cs.CV2025

GaussianOcc: Fully Self-supervised and Efficient 3D Occupancy Estimation with Gaussian Splatting

Wanshui Gan, Fang Liu, Hongbin Xu +2

We introduce GaussianOcc, a systematic method that investigates the two usages of Gaussian splatting for fully self-supervised and efficient 3D occupancy estimation in surround vie…

cs.CV2024

4DStyleGaussian: Zero-shot 4D Style Transfer with Gaussian Splatting

Wanlin Liang, Hongbin Xu, Weitao Chen +2

3D neural style transfer has gained significant attention for its potential to provide user-friendly stylization with spatial consistency. However, existing 3D style transfer metho…

cs.CV2025

LSVG: Language-Guided Scene Graphs with 2D-Assisted Multi-Modal Encoding for 3D Visual Grounding

Feng Xiao, Hongbin Xu, Guocan Zhao +1

3D visual grounding aims to localize the unique target described by natural languages in 3D scenes. The significant gap between 3D and language modalities makes it a notable challe…

cs.CV2025

StyleTailor: Towards Personalized Fashion Styling via Hierarchical Negative Feedback

Hongbo Ma, Fei Shen, Hongbin Xu +5

The advancement of intelligent agents has revolutionized problem-solving across diverse domains, yet solutions for personalized fashion styling remain underexplored, which holds im…

cs.CV2025

TextSplat: Text-Guided Semantic Fusion for Generalizable Gaussian Splatting

Zhicong Wu, Hongbin Xu, Gang Xu +6

Recent advancements in Generalizable Gaussian Splatting have enabled robust 3D reconstruction from sparse input views by utilizing feed-forward Gaussian Splatting models, achieving…

cs.CV2021

Self-supervised Multi-view Stereo via Effective Co-Segmentation and Data-Augmentation

Hongbin Xu, Zhipeng Zhou, Yu Qiao +2

Recent studies have witnessed that self-supervised methods based on view synthesis obtain clear progress on multi-view stereo (MVS). However, existing methods rely on the assumptio…

cs.CV2026

LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation

Jiazheng Xing, Fei Du, Hangjie Yuan +7

Recent advances in diffusion models have significantly improved text-to-video generation, enabling personalized content creation with fine-grained control over both foreground and…

cs.CV2022

CP-Net: Contour-Perturbed Reconstruction Network for Self-Supervised Point Cloud Learning

Mingye Xu, Yali Wang, Zhipeng Zhou +2

Self-supervised learning has not been fully explored for point cloud analysis. Current frameworks are mainly based on point cloud reconstruction. Given only 3D coordinates, such ap…

cs.CV2024

ControLRM: Fast and Controllable 3D Generation via Large Reconstruction Model

Hongbin Xu, Weitao Chen, Zhipeng Zhou +4

Despite recent advancements in 3D generation methods, achieving controllability still remains a challenging issue. Current approaches utilizing score-distillation sampling are hind…

eess.SP2026

Time-Frequency Mode Decomposition for Wind Turbine Vibration Monitoring under Variable Speed Operation

Wei Zhou, Wei-Jian Li, Desen Zhu +2

Wind turbine vibration monitoring under variable speed operation requires separating nonstationary rotor-order components whose frequencies and operating intervals depend on operat…

cs.CV2026

Yo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic Expansion

Keyang Lu, Sifan Zhou, Hongbin Xu +6

Realistic 3D city generation is fundamental to a wide range of applications, including virtual reality and digital twins. However, most existing methods rely on training a single d…

cs.CV2024

Improving 3D Finger Traits Recognition via Generalizable Neural Rendering

Hongbin Xu, Junduan Huang, Yuer Ma +2

3D biometric techniques on finger traits have become a new trend and have demonstrated a powerful ability for recognition and anti-counterfeiting. Existing methods follow an explic…

cs.CV2026

SemHiTok: A Unified Image Tokenizer via Semantic-Guided Hierarchical Codebook for Multimodal Understanding and Generation

Zisheng Chen, Chunwei Wang, Runhui Huang +6

In this paper, we introduce SemHiTok, a unified image Tokenizer via Semantic-Guided Hierarchical codebook that provides consistent discrete representations for multimodal understan…

cs.CV2024

StyleDyRF: Zero-shot 4D Style Transfer for Dynamic Neural Radiance Fields

Hongbin Xu, Weitao Chen, Feng Xiao +2

4D style transfer aims at transferring arbitrary visual style to the synthesized novel views of a dynamic 4D scene with varying viewpoints and times. Existing efforts on 3D style t…

physics.chem-ph2026

From Static Spectra to Operando Infrared Dynamics: Physics Informed Flow Modeling and a Benchmark

Shuquan Ye, Ben Fei, Hongbin Xu +2

The Solid Electrolyte Interphase (SEI) is critical to the performance of lithium-ion batteries, yet its analysis via Operando Infrared (IR) spectroscopy remains experimentally comp…

cs.RO2025

FedVLA: Federated Vision-Language-Action Learning with Dual Gating Mixture-of-Experts for Robotic Manipulation

Cui Miao, Tao Chang, Meihan Wu +4

Vision-language-action (VLA) models have significantly advanced robotic manipulation by enabling robots to interpret language instructions for task execution. However, training the…

cs.CV2025

PhysMamba: State Space Duality Model for Remote Physiological Measurement

Zhixin Yan, Yan Zhong, Hongbin Xu +4

Remote Photoplethysmography (rPPG) enables non-contact physiological signal extraction from facial videos, offering applications in psychological state analysis, medical assistance…

cs.CV2025

Cyc3D: Fine-grained Controllable 3D Generation via Cycle Consistency Regularization

Hongbin Xu, Chaohui Yu, Feng Xiao +5

Despite the remarkable progress of 3D generation, achieving controllability, i.e., ensuring consistency between generated 3D content and input conditions like edge and depth, remai…

cs.GR2025

GSsplat: Generalizable Semantic Gaussian Splatting for Novel-view Synthesis in 3D Scenes

Feng Xiao, Hongbin Xu, Wanlin Liang +1

The semantic synthesis of unseen scenes from multiple viewpoints is crucial for research in 3D scene understanding. Current methods are capable of rendering novel-view images and s…

cs.CV2024

SeCG: Semantic-Enhanced 3D Visual Grounding via Cross-modal Graph Attention

Feng Xiao, Hongbin Xu, Qiuxia Wu +1

3D visual grounding aims to automatically locate the 3D region of the specified object given the corresponding textual description. Existing works fail to distinguish similar objec…

cs.CV2023

CostFormer:Cost Transformer for Cost Aggregation in Multi-view Stereo

Weitao Chen, Hongbin Xu, Zhipeng Zhou +4

The core of Multi-view Stereo(MVS) is the matching process among reference and source pixels. Cost aggregation plays a significant role in this process, while previous methods focu…

cs.MM2025

Safe-VAR: Safe Visual Autoregressive Model for Text-to-Image Generative Watermarking

Ziyi Wang, Songbai Tan, Gang Xu +5

With the success of autoregressive learning in large language models, it has become a dominant approach for text-to-image generation, offering high efficiency and visual quality. H…

cs.CV2024

PointDC:Unsupervised Semantic Segmentation of 3D Point Clouds via Cross-modal Distillation and Super-Voxel Clustering

Zisheng Chen, Hongbin Xu, Weitao Chen +5

Semantic segmentation of point clouds usually requires exhausting efforts of human annotations, hence it attracts wide attention to the challenging topic of learning from unlabeled…

cs.CV2023

Artificial Intelligence System for Detection and Screening of Cardiac Abnormalities using Electrocardiogram Images

Deyun Zhang, Shijia Geng, Yang Zhou +28

The artificial intelligence (AI) system has achieved expert-level performance in electrocardiogram (ECG) signal analysis. However, in underdeveloped countries or regions where the…

cs.CV2023

Semi-supervised Deep Multi-view Stereo

Hongbin Xu, Weitao Chen, Yang Liu +5

Significant progress has been witnessed in learning-based Multi-view Stereo (MVS) under supervised and unsupervised settings. To combine their respective merits in accuracy and com…

cs.CV2026

UAM: A Dual-Stream Perspective on Forgetting in VLA Training

Jianke Zhang, Yuanfei Luo, Yucheng Hu +6

Vision--language--action (VLA) models are typically built by fine-tuning a pretrained vision--language model (VLM) on action data. However, we show that this standard recipe system…